| Title: | Base Class and Methods for 'gson' Format |
| Version: | 0.2.2 |
| Description: | Provides a lightweight container and exchange format for gene set collections. A 'GSON' object stores which genes belong to which gene set, together with gene set and gene names, the identifier types in use, species, versions and source metadata. A collection can be built from data frames, read from and written to the 'gson' JavaScript Object Notation (JSON) format and the 'GMT' format, subset by gene set, merged across sources, validated, and resolved to the web addresses of the databases it comes from, so that a collection gathered by one package can be analysed by another. |
| Imports: | jsonlite, methods, stats, utils, yulab.utils (≥ 0.0.7) |
| Suggests: | testthat (≥ 3.0.0) |
| ByteCompile: | true |
| License: | Artistic-2.0 |
| URL: | https://yulab-smu.top/biomedical-knowledge-mining-book/ |
| BugReports: | https://github.com/YuLab-SMU/gson/issues |
| Encoding: | UTF-8 |
| Config/testthat/edition: | 3 |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-10-09 12:10:29 UTC; wang |
| Author: | Guangchuang Yu |
| Maintainer: | Guangchuang Yu <guangchuangyu@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-09 12:40:07 UTC |
gson: Base Class and Methods for 'gson' Format
Description
Provides a lightweight container and exchange format for gene set collections. A 'GSON' object stores which genes belong to which gene set, together with gene set and gene names, the identifier types in use, species, versions and source metadata. A collection can be built from data frames, read from and written to the 'gson' JavaScript Object Notation (JSON) format and the 'GMT' format, subset by gene set, merged across sources, validated, and resolved to the web addresses of the databases it comes from, so that a collection gathered by one package can be analysed by another.
Author(s)
Maintainer: Guangchuang Yu guangchuangyu@gmail.com (ORCID) [copyright holder]
Authors:
Guangchuang Yu guangchuangyu@gmail.com (ORCID) [copyright holder]
See Also
Useful links:
Report bugs at https://github.com/YuLab-SMU/gson/issues
Class "GSON" This class represents gene set information.
Description
Class "GSON" This class represents gene set information.
Slots
gsid2genedata.frame with two columns of 'gsid' and 'gene'
gsid2namedata.frame with two columns of 'gsid' and 'name'
gene2namedata.frame with two columns of 'gene' and 'name'
schema_versionversion of the GSON file schema
speciesspecies of the annotation
gsnamegene set name, e.g., GO, KEGG
versionversion of the gene set
accessed_datetime to obtain the gene set data
keytypekeytype of genes
gsidtypeidentifier domain of the gene set IDs, e.g., GO, KEGG, REAC
urlpatternURL pattern to browse gene set online
infoextra information
Author(s)
Guangchuang Yu https://yulab-smu.top
Subset a GSON object
Description
x[i] keeps the gene sets selected by i and returns a GSON, with the
three mapping tables pruned together: gene2name is filtered by the genes
that are still a member of a kept gene set, so the subset cannot grow a
reference to a gene that is no longer there. The metadata slots carry over
unchanged. x[[i]] returns the genes of one gene set.
Usage
## S3 method for class 'GSON'
x[i, j, ..., drop = FALSE]
## S3 method for class 'GSON'
x[[i, ...]]
Arguments
x |
A |
i |
Gene set IDs, or a logical or numeric index into |
j |
Unused. Supplying a second index is an error; a |
... |
Unused. |
drop |
Ignored; a |
Value
[.GSON returns a GSON, [[.GSON a character vector of genes.
Coerce GSON to a data frame
Description
Coerce GSON to a data frame
Usage
## S3 method for class 'GSON'
as.data.frame(x, row.names = NULL, optional = FALSE, ...)
Arguments
x |
A |
row.names |
Unused. |
optional |
Unused. |
... |
Unused. |
Value
A data frame with gene set-gene memberships and optional names.
construct a 'GSON' object
Description
construct a 'GSON' object
Usage
gson(
gsid2gene,
gsid2name = NULL,
gene2name = NULL,
schema_version = "1.0",
species = NULL,
gsname = NULL,
version = NULL,
accessed_date = NULL,
keytype = NULL,
gsidtype = NULL,
urlpattern = NULL,
info = NULL
)
Arguments
gsid2gene |
A data frame with first column of gene set IDs and second column of genes |
gsid2name |
A data frame with first column of gene set IDs and second column of gene set names |
gene2name |
A data frame with first column of genes and second column of gene symbols |
schema_version |
GSON file schema version |
species |
Which species of the genes belongs to |
gsname |
Name of the gene set (e.g., GO, KEGG, etc.) |
version |
version of the gene set |
accessed_date |
date to obtain the gene set data |
keytype |
keytype of genes |
gsidtype |
identifier domain of the gene set IDs, e.g., GO, KEGG, REAC |
urlpattern |
URL pattern |
info |
extra information |
Value
A 'GSON' instance
Examples
wpfile <- system.file('extdata', "wikipathways-20220310-gmt-Homo_sapiens.gmt", package='gson')
x <- read.gmt.wp(wpfile)
gsid2gene <- data.frame(gsid=x$wpid, gene=x$gene)
gsid2name <- unique(data.frame(gsid=x$wpid, name=x$name))
species <- unique(x$species)
version <- unique(x$version)
gson(gsid2gene=gsid2gene, gsid2name=gsid2name, species=species, version=version)
construct a 'GSONList' object
Description
construct a 'GSONList' object
Usage
gsonList(...)
Arguments
... |
input GSON objects |
Value
A 'GSONList' instance
Combine gene set collections
Description
gson_union() merges GSON objects into one collection. Two objects are the
same source when they declare the same gsname, and – once every object
declares one – the same gsidtype; the gene sets of one source are joined by
ID, so a gene set that appears in two files of one collection becomes one set
holding every gene either file gives it. A gene set ID that two objects from
different sources both use is a conflict, because the merged membership table
would say that one ID names two sets of genes, and .conflict decides what
happens: "error" refuses, "prefix" qualifies every gene set ID with the
source it comes from (KEGG:hsa00010), and "rename" qualifies only the IDs
that actually collide, leaving the rest of the collection alone. Where two
objects name the same gene set or the same gene differently, the first object
that says something wins.
Usage
gson_union(x = NULL, y = NULL, ..., .conflict = c("error", "prefix", "rename"))
Arguments
x, y |
A |
... |
More |
.conflict |
What to do about a gene set ID used by two different
sources: |
Details
The metadata is merged by what the merged object can still honestly say:
species, keytype and schema_version have to agree or the call is an
error, because a membership table holds genes of one organism identified by
one identifier type, while gsname, gsidtype, version, accessed_date
and info are joined with " + " when the sources name themselves
differently. urlpattern is kept only when every object declares the same
one and no gene set ID was renamed – a pattern that resolves to a wrong page
for part of the collection is worse than no pattern, and gson_url() says so.
Value
A GSON object holding every gene set of every input.
Examples
wpfile <- system.file('extdata', "wikipathways-20220310-gmt-Homo_sapiens.gmt", package='gson')
wp <- read.gmt.wp(wpfile, output = "GSON")
gson_union(wp[1:2], wp[3:4])
URLs of the gene sets of a GSON
Description
gson_url() substitutes each gene set ID into the {gsid} token of
x@urlpattern, and browseGS() opens the resulting URL in a browser. The
URL only comes from the object itself: gson ships no table of other
databases' websites, so a producer that wants its gene sets to be browsable
has to declare urlpattern.
Usage
gson_url(x, gsid = NULL)
browseGS(x, gsid, ...)
Arguments
x |
A |
gsid |
Gene set IDs to build URLs for. Defaults to |
... |
Further arguments passed to |
Value
gson_url() a named character vector of URLs, every entry
NA_character_ and a warning when the object carries no urlpattern.
browseGS() returns NULL invisibly.
Examples
wpfile <- system.file('extdata', "wikipathways-20220310-gmt-Homo_sapiens.gmt", package='gson')
wp <- read.gmt.wp(wpfile, output = "GSON")
gson_url(wp)[1:2]
gson_url(wp, c("WP100", "WP106"))
# browseGS(wp, "WP100") opens the same URL in a browser
read.gmt
Description
parse gmt file to a data.frame
write a GSON object to GMT format
Usage
read.gmt(gmtfile)
read.gmt.wp(gmtfile, output = "data.frame")
write.gmt(x, file = "")
Arguments
gmtfile |
gmt file |
output |
one of 'data.frame' or 'GSON' |
x |
A |
file |
output GMT file |
Value
data.frame
Author(s)
Guangchuang Yu
read and write gson file
Description
read and write gson file
Usage
read.gson(file)
write.gson(x, file = "")
Arguments
file |
A gson file |
x |
A |
Value
A GSON instance
Examples
wpfile <- system.file('extdata', "wikipathways-20220310-gmt-Homo_sapiens.gmt", package='gson')
x <- read.gmt.wp(wpfile, output = "GSON")
f = tempfile(fileext = '.gson')
write.gson(x, f)
read.gson(f)
show method
Description
show method for GSON instance
Usage
show(object)
Arguments
object |
A |
Value
message
Author(s)
Guangchuang Yu https://yulab-smu.top
Validate a GSON object
Description
validate_gson() checks the core data contract of a GSON gene set
collection. It is useful after constructing an object manually or reading one
from an external source.
Usage
validate_gson(x, error = TRUE)
Arguments
x |
A |
error |
Logical. If |
Details
Nothing is rewritten. A value the object declares is reported as it stands –
a schema_version this package does not know how to read, a keytype or
gsidtype that names no identifier type, a gsidtype that the gene set IDs
of the object itself contradict – because a reader that knows what a file
says is not the one that decides what the producer meant.
Value
TRUE if the object is valid. If error = FALSE, returns a
character vector of validation messages when invalid.
Examples
x <- gson(data.frame(gsid = "GS1", gene = "gene1"))
validate_gson(x)