Contents

1 Introduction

ConvergeR is an analytical toolkit designed to compute, visualize, and statistically evaluate the convergence of single-cell states, based on the methodology established by Whitfield et al. (2026).

In this vignette, we will demonstrate the package workflow using a full dataset of 3,000 Peripheral Blood Mononuclear Cells (PBMC). We will evaluate cells for their balance between two opposing states: Immune Signaling UP versus Immune Signaling DN from MSigDB.

2 1. Environment Setup and Data Preparation

Let’s load the required packages and the lightweight subset of the PBMC dataset bundled with the package.

library(ConvergeR)
library(Seurat)
#> Le chargement a nécessité le package : SeuratObject
#> Le chargement a nécessité le package : sp
#> 
#> Attachement du package : 'SeuratObject'
#> Les objets suivants sont masqués depuis 'package:base':
#> 
#>     intersect, t
library(ggplot2)

# Load the built-in example dataset
seu_path <- system.file("extdata", "pbmc3k_subset.rds", package = "ConvergeR")
seu <- readRDS(seu_path)

# Normalising
seu <- NormalizeData(seu, verbose = FALSE)

# Inject mock clinical metadata to simulate a multi-patient cohort
set.seed(42)
seu$age <- sample(20:70, ncol(seu), replace = TRUE)
seu$tumor_status <- sample(c("Primary", "Metastatic"), ncol(seu), replace = TRUE)
seu$treatment <- sample(c("Treated", "Untreated"), ncol(seu), replace = TRUE)
seu$SS <- paste0("Patient_", sample(1:5, ncol(seu), replace = TRUE))

3 2. Importing MSigDB Signatures and Scoring

Instead of manually importing .gmt files, we use the msigdbr package to directly retrieve the official Hallmark Immune signatures.

library(msigdbr)

# S'assurer que l'on travaille sur l'assay RNA principal
DefaultAssay(seu) <- "RNA"

# Récupérer deux signatures immunitaires/inflammatoires bien exprimées dans les PBMC
hallmark_sets <- msigdbr(species = "Homo sapiens", category = "H")

raw_group1 <- hallmark_sets[hallmark_sets$gs_name == "HALLMARK_INTERFERON_ALPHA_RESPONSE", ]$gene_symbol
raw_group2 <- hallmark_sets[hallmark_sets$gs_name == "HALLMARK_INFLAMMATORY_RESPONSE", ]$gene_symbol

# Intersection stricte avec les gènes réellement présents dans pbmc3k_subset
group1_genes <- intersect(raw_group1, rownames(seu))
group2_genes <- intersect(raw_group2, rownames(seu))

# Calcul des scores de modules Seurat
seu <- AddModuleScore(seu, features = list(group1_genes), name = "Score_IFN_")
seu <- AddModuleScore(seu, features = list(group2_genes), name = "Score_Inflam_")

# Renommer proprement les colonnes pour la suite
seu$Score_IFN <- seu$Score_IFN_1
seu$Score_Inflam <- seu$Score_Inflam_1

4 3. Computing the Convergence Score

seu <- CalculateConvergedScore(
  seurat_obj = seu,
  principal_score = "Score_IFN",
  other_scores = "Score_Inflam",
  principal_name = "Interferon",
  other_name = "Inflammatory",
  output_colname = "Converged_Immune",
  principal_color = "#9b59b6", # Violet pour Interféron
  other_color = "#e67e22"      # Orange pour Inflammatoire
)

5 4. Visualization

ConvergeR automatically reads the color metadata generated in the previous step to produce standardized plots.

5.1 Proportion Barplot

Proportion of cells converging to either Immune state across our 10 mock patients.

PlotConvergenceProportion(
  seurat_obj = seu, 
  x_var = "SS", 
  fill_var = "Converged_Immune_Direction", 
  title = "Immune Convergence by Patient"
)

5.2 Crossed Proportion Barplot

Faceted by tumor status to identify subgroup trends.

PlotConvergenceCrossedProportion(
  seurat_obj = seu, 
  x_var = "SS",
  fill_var = "Converged_Immune_Direction",
  facet_var = "tumor_status",
  title = "Immune Convergence split by Tumor Status"
)

5.3 Density Distribution

Continuous relationship of the convergence direction with patient age.

PlotConvergenceDensity(
  seurat_obj = seu, 
  x_var = "age",
  fill_var = "Converged_Immune_Direction",
  x_label = "Patient Age"
)

6 5. Statistical Testing

We evaluate the statistical significance at the patient level (pseudobulk) to prevent pseudoreplication bias caused by single-cell testing.

6.1 Bivariate Test

TestConvergenceScore(
  seurat_obj = seu, 
  score_var = "Converged_Immune", 
  test_var = "age",
  level = "patient", 
  patient_id_var = "SS"
)
#> Spearman correlation | level: patient (pseudobulk) | N = 5 | testing 'Converged_Immune' by 'age' | p = 0.517

How to read this result: The output message displays the statistical test used (e.g., Spearman correlation for continuous variables), the analysis level (patient pseudobulk to avoid single-cell pseudoreplication), the number of observations (\(N\) patients), and the p-value. A p-value below 0.05 indicates a statistically significant association between the convergence score and the tested variable.

6.2 Multivariate Linear Model

TestMultivariateConvergence(
  seurat_obj = seu, 
  score_var = "Converged_Immune", 
  test_vars = c("age", "tumor_status", "treatment"),
  level = "patient", 
  patient_id_var = "SS"
)
#> ======================================================
#> Multivariate Linear Model | level: patient (pseudobulk) | N = 5
#> Formula: Converged_Immune ~ age + tumor_status + treatment
#> Global Model R-squared: 0.77 | Global p-value: 0.586
#> ------------------------------------------------------
#> Independent Variable Effects:
#>   - age                       : p = 0.563     (Estimate =  0.002)
#>   - tumor_statusPrimary       : p = 0.599     (Estimate = -0.079)
#>   - treatmentUntreated        : p = 0.551     (Estimate =  0.078)
#> ======================================================

How to read this result: This model fits a multivariable linear regression on pseudobulk data. It details the global model performance (\(R^2\) and global p-value) and breaks down the independent effect (Estimate) and significance (p-value accompanied by conventional significance stars) for each covariate, allowing you to assess the effect of a variable while controlling for potential confounders.

7 Session Information

sessionInfo()
#> R version 4.4.3 (2025-02-28)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 22.04.5 LTS
#> 
#> Matrix products: default
#> BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.20.so;  LAPACK version 3.10.0
#> 
#> locale:
#>  [1] LC_CTYPE=fr_FR.UTF-8       LC_NUMERIC=C              
#>  [3] LC_TIME=fr_FR.UTF-8        LC_COLLATE=C              
#>  [5] LC_MONETARY=fr_FR.UTF-8    LC_MESSAGES=fr_FR.UTF-8   
#>  [7] LC_PAPER=fr_FR.UTF-8       LC_NAME=C                 
#>  [9] LC_ADDRESS=C               LC_TELEPHONE=C            
#> [11] LC_MEASUREMENT=fr_FR.UTF-8 LC_IDENTIFICATION=C       
#> 
#> time zone: Europe/Paris
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats     graphics  grDevices utils     datasets  methods   base     
#> 
#> other attached packages:
#> [1] msigdbr_26.1.1     ggplot2_4.0.3      Seurat_5.1.0       SeuratObject_5.1.0
#> [5] sp_2.2-3           ConvergeR_0.99.1   BiocStyle_2.32.1  
#> 
#> loaded via a namespace (and not attached):
#>   [1] deldir_2.0-4           pbapply_1.7-5          gridExtra_2.3.1       
#>   [4] rlang_1.3.0            magrittr_2.0.5         RcppAnnoy_0.0.23      
#>   [7] otel_0.2.0             spatstat.geom_3.8-2    matrixStats_1.5.0     
#>  [10] ggridges_0.5.7         compiler_4.4.3         png_0.1-9             
#>  [13] vctrs_0.7.3            reshape2_1.4.5         stringr_1.6.0         
#>  [16] pkgconfig_2.0.3        fastmap_1.2.0          magick_2.9.1          
#>  [19] labeling_0.4.3         promises_1.5.0         rmarkdown_2.32        
#>  [22] tinytex_0.60           purrr_1.2.2            xfun_0.60             
#>  [25] cachem_1.1.0           jsonlite_2.0.0         goftest_1.2-3         
#>  [28] later_1.4.8            spatstat.utils_3.2-4   irlba_2.3.7           
#>  [31] parallel_4.4.3         cluster_2.1.8.3        R6_2.6.1              
#>  [34] ica_1.0-3              spatstat.data_3.1-9    stringi_1.8.9         
#>  [37] bslib_0.12.0           RColorBrewer_1.1-3     reticulate_1.47.0     
#>  [40] spatstat.univar_3.2-0  parallelly_1.48.0      lmtest_0.9-40         
#>  [43] jquerylib_0.1.4        scattermore_1.2        assertthat_0.2.1      
#>  [46] Rcpp_1.1.2             bookdown_0.48          knitr_1.52            
#>  [49] tensor_1.5.1           future.apply_1.20.2    zoo_1.9-0             
#>  [52] sctransform_0.4.3      httpuv_1.6.17          Matrix_1.7-3          
#>  [55] splines_4.4.3          igraph_2.3.3           tidyselect_1.2.1      
#>  [58] abind_1.4-8            rstudioapi_0.19.0      dichromat_2.0-1       
#>  [61] yaml_2.3.12            spatstat.random_3.5-1  spatstat.explore_3.8-2
#>  [64] codetools_0.2-19       miniUI_0.1.2           listenv_1.0.0         
#>  [67] plyr_1.8.9             lattice_0.22-5         tibble_3.3.1          
#>  [70] withr_3.0.3            shiny_1.14.0           S7_0.2.2              
#>  [73] ROCR_1.0-12            evaluate_1.0.5         Rtsne_0.17            
#>  [76] future_1.75.0          fastDummies_1.7.6      survival_3.8-3        
#>  [79] polyclip_1.10-7        fitdistrplus_1.2-6     pillar_1.11.1         
#>  [82] BiocManager_1.30.27    KernSmooth_2.23-26     plotly_4.12.1         
#>  [85] generics_0.1.4         RcppHNSW_0.7.0         scales_1.4.0          
#>  [88] globals_0.19.1         xtable_1.8-8           glue_1.8.1            
#>  [91] tools_4.4.3            data.table_1.18.6.1    RSpectra_0.16-2       
#>  [94] RANN_2.6.3             leiden_0.4.3.1         dotCall64_1.2         
#>  [97] cowplot_1.2.0          grid_4.4.3             tidyr_1.3.2           
#> [100] nlme_3.1-168           patchwork_1.3.2        cli_3.6.6             
#> [103] spatstat.sparse_3.2-0  spam_2.11-4            viridisLite_0.4.3     
#> [106] dplyr_1.2.1            uwot_0.2.5             gtable_0.3.6          
#> [109] sass_0.4.10            digest_0.6.39          progressr_1.0.0       
#> [112] ggrepel_0.9.8          htmlwidgets_1.6.4      farver_2.1.2          
#> [115] htmltools_0.5.9        lifecycle_1.0.5        httr_1.4.9            
#> [118] mime_0.13              MASS_7.3-65