Provides a lightweight container and exchange format for gene set collections. A ‘GSON’ object stores which genes belong to which gene set, together with gene set and gene names, the identifier types in use, species, versions and source metadata. A collection can be built from data frames, read from and written to the ‘gson’ JavaScript Object Notation (JSON) format and the ‘GMT’ format, subset by gene set, merged across sources, validated, and resolved to the web addresses of the databases it comes from, so that a collection gathered by one package can be analysed by another.
A gene set collection answers one question: which genes belong to
which gene set. GSON holds that answer together with the
metadata needed to read it: what the identifiers are, which organism,
which release, where the sets came from. The file it writes is plain
JSON, so a collection can be read outside R too.
# the released version, from CRAN
install.packages("gson")
# the development version, from GitHub
# install.packages("remotes")
remotes::install_github("YuLab-SMU/gson")Three tables make up a GSON. gsid2gene is
required and holds one row per gene set and gene. gsid2name
and gene2name are optional and map gene set IDs and gene
IDs to readable names. The rest is metadata: keytype names
the identifier type of the genes, gsidtype the identifier
domain of the gene set IDs, and species,
gsname, version, accessed_date,
urlpattern and info say where the collection
came from.
library(gson)
gsid2gene <- data.frame(
gsid = c("GO:0007049", "GO:0007049", "GO:0006915"),
gene = c("101", "102", "103")
)
gsid2name <- data.frame(
gsid = c("GO:0007049", "GO:0006915"),
name = c("cell cycle", "apoptotic process")
)
x <- gson(
gsid2gene,
gsid2name = gsid2name,
species = "Homo sapiens",
gsname = "example",
version = "2026-06-30",
keytype = "ENTREZID",
gsidtype = "GO"
)
x
#> >> Gene Set: example
#> >> 3 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30
#> >> Gene set ID type: GOgson() keeps the columns of the three tables character.
A factor would retain the levels it no longer uses, and a collection is
read by split()-ing it by gene set ID, so one unused level
becomes one empty gene set that length(),
names() and both writers then report as real.
f <- tempfile(fileext = ".gson")
write.gson(x, f)
cat(readLines(f), sep = "\n")
#> {
#> "gsid2gene": {
#> "GO:0007049": ["101", "102"],
#> "GO:0006915": ["103"]
#> },
#> "gsid2name": {
#> "gsid": ["GO:0007049", "GO:0006915"],
#> "name": ["cell cycle", "apoptotic process"]
#> },
#> "gene2name": null,
#> "schema_version": ["1.0"],
#> "species": ["Homo sapiens"],
#> "gsname": ["example"],
#> "version": ["2026-06-30"],
#> "accessed_date": null,
#> "keytype": ["ENTREZID"],
#> "gsidtype": ["GO"],
#> "urlpattern": null,
#> "info": null
#> }y <- read.gson(f)
identical(x, y)
#> [1] TRUEread.gson() opens a .gson, a
.gson.gz, and a file written by any other program that
follows the format. A file it cannot make sense of is answered by
name:
read.gson("no-such-file.gson")
#> Error:
#> ! the gson file no-such-file.gson does not existA file that exists but is not a collection is answered by the member
that is wrong and the shape it holds.
tests/testthat/test-io.R covers the shapes a hand-written
.gson turns up with.
GMT is the other side of the same object. write.gmt()
gives one line per gene set, holding the ID, a description and the
genes; read.gmt() reads a plain GMT back as a two-column
table.
gmt <- tempfile(fileext = ".gmt")
write.gmt(x, gmt)
cat(readLines(gmt), sep = "\n")
#> GO:0006915 apoptotic process 103
#> GO:0007049 cell cycle 101 102A membership row with no gene set ID or no gene cannot go into either
format, so both writers drop those rows and say how many.
validate_gson() reports them instead, on an object you can
still look at.
wpfile <- system.file(
"extdata",
"wikipathways-20220310-gmt-Homo_sapiens.gmt",
package = "gson"
)
wp <- read.gmt.wp(wpfile, output = "GSON")
wp
#> >> Gene Set: WikiPathways
#> >> 7781 genes annotated by 724 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WPread.gmt.wp() is the WikiPathways reader. It takes the
species and the release out of the file name, records
gsidtype = "WP", and sets urlpattern to the
address WikiPathways serves its pathways from. Your own GMT follows none
of those conventions, so read it with read.gmt() and build
the object:
term2gene <- read.gmt(gmt)
head(term2gene)
#> term gene
#> 1 GO:0006915 103
#> 2 GO:0007049 101
#> 3 GO:0007049 102
from_gmt <- gson(
term2gene,
species = "Homo sapiens",
gsname = "my sets",
version = "1.0",
keytype = "ENTREZID"
)
from_gmt
#> >> Gene Set: my sets
#> >> 3 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 1.0gson() renames the first two columns of the table it is
given to gsid and gene, so the
term column of read.gmt() becomes the gene set
ID here.
gson_url(wp, c("WP100", "WP106"))
#> WP100
#> "https://www.wikipathways.org/pathways/WP100.html"
#> WP106
#> "https://www.wikipathways.org/pathways/WP106.html"gson_url() substitutes a gene set ID for the
{gsid} token of the pattern the object declares. With no
gsid it covers every gene set, in the order
names(x) holds them, and browseGS(wp, "WP100")
opens the same address in the browser. gson ships no table of database
websites, since it has no way to check one, so an address comes only
from the object:
mine <- gson(gsid2gene, gsname = "my sets", gsidtype = "GO")
gson_url(mine)
#> Warning: this GSON carries no 'urlpattern', so its gene sets have no URL
#> GO:0007049 GO:0006915
#> NA NA
browseGS(mine, "GO:0007049")
#> Error:
#> ! this GSON carries no 'urlpattern', so its gene sets have no URL to browsex[i] keeps the gene sets i selects and
returns a GSON. A filtered collection stays a collection,
rather than three tables you filtered by hand and kept in step
yourself.
wp2 <- wp[c("WP100", "WP106")]
wp2
#> >> Gene Set: WikiPathways
#> >> 35 genes annotated by 2 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WPGene set IDs, positions and logical vectors all select,
x[] returns x, and x[[i]] still
returns the genes of one gene set. The three tables are pruned together,
so gene2name loses any gene that is no longer a member of
something kept, and every metadata slot carries over to the result. Most
subsets come from the gene set names:
keep <- wp@gsid2name[grepl("apoptosis", wp@gsid2name$name, ignore.case = TRUE), ]
wp[keep$gsid]
#> >> Gene Set: WikiPathways
#> >> 211 genes annotated by 9 gene sets.
#> >> Species: Homo sapiens
#> >> Version: WikiPathways_20220310
#> >> Gene set ID type: WPA GSON has no second dimension, so wp[, 2]
is an error, and so is an index out of bounds or a gene set ID the
object does not hold.
gson_union() merges several GSON objects,
or one GSONList, into a single collection. Two objects that
name the same source are two parts of one collection, so a gene set both
carry joins into one set holding every gene either gives it:
monday <- gson(
data.frame(gsid = c("GO:0007049", "GO:0006915"), gene = c("101", "103")),
gsid2name = data.frame(gsid = c("GO:0007049", "GO:0006915"),
name = c("cell cycle", "apoptotic process")),
gsname = "example", gsidtype = "GO", species = "Homo sapiens",
version = "2026-06-30", keytype = "ENTREZID"
)
tuesday <- gson(
data.frame(gsid = c("GO:0006915", "GO:0008150"), gene = c("103", "104")),
gsname = "example", gsidtype = "GO", species = "Homo sapiens",
version = "2026-06-30", keytype = "ENTREZID"
)
merged <- gson_union(monday, tuesday)
merged
#> >> Gene Set: example
#> >> 3 genes annotated by 3 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30
#> >> Gene set ID type: GO
merged@gsid2gene
#> gsid gene
#> 1 GO:0007049 101
#> 2 GO:0006915 103
#> 3 GO:0008150 104Two objects from different sources that both use the same gene set ID are refused:
custom <- gson(
data.frame(gsid = "GO:0007049", gene = "999"),
gsname = "my sets", gsidtype = "GO", species = "Homo sapiens",
version = "1.0", keytype = "ENTREZID"
)
gson_union(monday, custom)
#> Error:
#> ! gene set ID(s) used by more than one source: GO:0007049 -- merge with .conflict = "prefix" to qualify every ID with its source, or .conflict = "rename" to qualify only theseThe merged table would say that one ID names two sets of genes, and a
result read from it would be wrong without looking wrong. Ask for what
you want instead: .conflict = "rename" qualifies only the
colliding IDs, .conflict = "prefix" qualifies all of them,
and the gsid2name rows follow the IDs they describe.
merged <- gson_union(monday, custom, .conflict = "rename")
merged
#> >> Gene Set: example + my sets
#> >> 3 genes annotated by 3 gene sets.
#> >> Species: Homo sapiens
#> >> Version: 2026-06-30 + 1.0
#> >> Gene set ID type: GO
names(merged)
#> [1] "example:GO:0007049" "GO:0006915" "my sets:GO:0007049"species, keytype and
schema_version have to agree between the objects, because
one membership table cannot mean two organisms, two gene identifier
types or two file schemas. gsname, gsidtype,
version, accessed_date and info
are joined with + when the sources name themselves
differently. urlpattern is kept only when every object
declares the same one and no ID moved, because a pattern that resolves
part of the collection to the wrong page is worse than no pattern.
A GSONList is a list of collections.
gson_union() takes one as its input, and x[1]
of a GSONList is still a GSONList.
gson() refuses what the class cannot hold: a table that
is not a data frame, one that lacks the documented columns, identifiers
that are not character. validate_gson() reports what a
valid object can still get wrong. With error = FALSE it
returns the report as a character vector instead of throwing it.
validate_gson(x)
#> [1] TRUE
broken <- gson(
data.frame(gsid = c("GO:0007049", "GO:0007049"), gene = c("101", "101")),
gsid2name = data.frame(gsid = "GO:0007049", name = "cell cycle"),
gsidtype = "KEGG", keytype = "UNKNOWN", species = "Homo sapiens"
)
validate_gson(broken, error = FALSE)
#> [1] "gsid2gene contains duplicate gsid-gene memberships"
#> [2] "keytype 'UNKNOWN' names no identifier type; leave it NULL when the type is not known"
#> [3] "gsidtype 'KEGG' is not the prefix of the gene set IDs, which begin GO -- as in 'GO:0007049'"There is deliberately no dictionary of allowed keytype
or gsidtype values. Producers in this family write
kegg_orthology, eggNOG_OG and
hmdb_id, none of them a value any annotation package
offers, and an enum held by the format would reject collections that
work perfectly well. What gson can verify is an object against itself,
so a gsidtype is compared with the prefixes the gene set
IDs of that object actually carry, and nothing is ever rewritten.
enrichit takes the object as it is;
ora_gson(), gsea_gson() and
nsea_gson() all have a gson argument that
wants a GSON.
geneList <- sort(setNames(rnorm(30), as.character(101:130)), decreasing = TRUE)
gsea <- enrichit::gsea_gson(geneList, gson = x)
ora <- enrichit::ora_gson(gene = c("101", "102"), gson = x,
pvalueCutoff = 1, minGSSize = 1)A tool that wants the bare tables finds them in the slots.
gsid2name is the map from gene set ID to label that puts a
readable name on a result, and keytype is what tells an
annotation package how to read the genes.
head(x@gsid2gene)
#> gsid gene
#> 1 GO:0007049 101
#> 2 GO:0007049 102
#> 3 GO:0006915 103
head(x@gsid2name)
#> gsid name
#> 1 GO:0007049 cell cycle
#> 2 GO:0006915 apoptotic process{
"gsid2gene": {"GO:0007049": ["101", "102"], "GO:0006915": ["103"]},
"gsid2name": {"gsid": ["GO:0007049", "GO:0006915"],
"name": ["cell cycle", "apoptotic process"]},
"gene2name": null,
"schema_version": ["1.0"],
"species": ["Homo sapiens"],
"gsname": ["example"],
"version": ["2026-06-30"],
"accessed_date": null,
"keytype": ["ENTREZID"],
"gsidtype": ["GO"],
"urlpattern": null,
"info": null
}gsid2gene is required. It is an object keyed by gene
set ID, each key holding an array of gene identifiers.gsid2name and gene2name are optional. Each
is an object of two arrays, gsid/gene and
name.schema_version describes the file schema, currently
1.0. It is not the package version. A file whose version
gson does not know is read anyway and reported by
validate_gson().gsidtype names the domain the gene set IDs are drawn
from, GO, KEGG or REAC for
instance, and is independent of keytype, which describes
the genes. A file written before the field existed reads back with
gsidtype set to NULL.species, gsname, version,
accessed_date, keytype,
urlpattern and info describe the source and
the interpretation of the collection.Because gsid2gene is keyed by gene set ID, reading a
file back groups the memberships by gsid. The gene sets
keep the order they were written in, and so do the genes inside each
set; the row order of a table that was not already grouped by gene set
ID does not survive. Row names do not survive either, and JSON has no
factor type, which is why gson() holds the columns as
character.
Guangchuang YU https://yulab-smu.top
School of Basic Medical Sciences, Southern Medical University
We welcome any contributions! By participating in this project you
agree to follow the conventions the rest of the
clusterProfiler family uses. Please open an issue at https://github.com/YuLab-SMU/gson/issues before a large
change.