| Title: | Exploratory Subgroup Identification in Survival and GLM Outcomes |
| Version: | 0.4.0 |
| Description: | Implements statistical methods for exploratory subgroup identification in clinical trials. Provides tools for identifying patient subgroups with differential treatment effects using machine learning approaches including Generalized Random Forests (GRF), LASSO regularization, and exhaustive combinatorial search algorithms. Supports survival endpoints (Cox proportional hazards), binary outcomes (log odds ratio, log relative risk, risk difference), continuous outcomes (mean difference), and count / rate outcomes (log incidence rate ratio via Poisson, quasi-Poisson, or negative-binomial GLMs with optional person-time offset). Features bootstrap bias correction using infinitesimal jackknife methods to address selection bias in post-hoc analyses. Designed for clinical researchers conducting exploratory subgroup analyses in randomized controlled trials, particularly for multi-regional clinical trials (MRCT) requiring regional consistency evaluation. Methods are described in Leon et al. (2024) <doi:10.1002/sim.10163>. Post-selection inference for the identified subgroup follows León and Anderson (2026) <doi:10.48550/arXiv.2609.38361>. |
| License: | MIT + file LICENSE |
| Depends: | R (≥ 4.1.0) |
| Encoding: | UTF-8 |
| VignetteBuilder: | knitr, quarto |
| SystemRequirements: | Quarto command line tools (https://github.com/quarto-dev/quarto-cli) |
| Imports: | data.table, doFuture, dplyr, foreach, future, future.apply, future.callr, ggplot2, glmnet (≥ 5.0), grf, gt, Matrix, mvtnorm, parallel, patchwork, policytree, progressr, randomForest, rlang, stats, survival, utils, weightedsurv (≥ 0.1.0) |
| Suggests: | DiagrammeR, doRNG, htmltools, MASS, sandwich, tidyr, forestploter, cubature, svglite, knitr, rmarkdown, katex, latentcor, quarto, testthat (≥ 3.0.0) |
| URL: | https://github.com/larry-leon/forestsearch, https://larry-leon.github.io/forestsearch/ |
| BugReports: | https://github.com/larry-leon/forestsearch/issues |
| Config/testthat/edition: | 3 |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-10-10 04:12:05 UTC; larryleon |
| Author: | Larry Leon [aut, cre] |
| Maintainer: | Larry Leon <larry.leon.05@post.harvard.edu> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-10 04:40:02 UTC |
forestsearch: Exploratory Subgroup Identification in Survival and GLM Outcomes
Description
Implements statistical methods for exploratory subgroup identification in clinical trials. Provides tools for identifying patient subgroups with differential treatment effects using machine learning approaches including Generalized Random Forests (GRF), LASSO regularization, and exhaustive combinatorial search algorithms. Supports survival endpoints (Cox proportional hazards), binary outcomes (log odds ratio, log relative risk, risk difference), continuous outcomes (mean difference), and count / rate outcomes (log incidence rate ratio via Poisson, quasi-Poisson, or negative-binomial GLMs with optional person-time offset). Features bootstrap bias correction using infinitesimal jackknife methods to address selection bias in post-hoc analyses. Designed for clinical researchers conducting exploratory subgroup analyses in randomized controlled trials, particularly for multi-regional clinical trials (MRCT) requiring regional consistency evaluation. Methods are described in Leon et al. (2024) doi:10.1002/sim.10163. Post-selection inference for the identified subgroup follows León and Anderson (2026) doi:10.48550/arXiv.2609.38361.
Author(s)
Maintainer: Larry Leon larry.leon.05@post.harvard.edu
Authors:
Larry Leon larry.leon.05@post.harvard.edu
See Also
Useful links:
Report bugs at https://github.com/larry-leon/forestsearch/issues
Normalise adjustment terms (mirrors forestsearch:::.fs_adjust_terms)
Description
Normalise adjustment terms (mirrors forestsearch:::.fs_adjust_terms)
Usage
.consistency_adj_terms(adjust_covariates)
Per-subject Cox pieces for the consistency approximation
Description
Fits the treatment-only (or covariate-adjusted) Cox model once on a subgroup
and returns the treatment log hazard ratio, the per-subject treatment
dfbeta influence contributions, and the split-perturbation SD.
Usage
.consistency_cox_pieces(
df,
tte.name = "Y",
event.name = "Event",
treat.name = "Treat",
cox_init = 0,
adjust_covariates = NULL
)
Arguments
df |
data.frame/data.table with the survival columns (and any columns
referenced by |
tte.name, event.name, treat.name |
Character; time, event, and (0/1)
treatment column names (defaults |
cox_init |
Numeric; warm-start for the unadjusted Cox coefficient. |
adjust_covariates |
Character vector or |
Value
List with beta_hat, dfbeta, sigma_D, n, d, log_scale
(TRUE), measure ("HR"); or NULL if degenerate.
Per-subject GLM pieces for the consistency approximation
Description
Fits an unweighted, unadjusted treatment-effect model and returns the
treatment coefficient, its per-subject dfbeta influence contributions, and
the split-perturbation SD.
Usage
.consistency_glm_pieces(
df,
outcome_type,
effect_measure = NULL,
treat.name = "Treat",
outcome.name = "Y",
offset.name = NULL,
adjust_covariates = NULL,
adverse_outcome = TRUE
)
Arguments
df |
data.frame/data.table with |
outcome_type |
Character; |
effect_measure |
Character or |
treat.name, outcome.name, offset.name |
Character; column names. |
adjust_covariates |
Character vector or |
adverse_outcome |
Logical; if |
Details
Supported: "OR" (logistic), "RR" (log-binomial with Poisson fallback),
"RD" (identity-link binomial), "MD" (OLS), "IRR" (Poisson with
offset(log(time))). "IRD" is not a single coefficient and returns
NULL, and the caller falls back to splitting.
Value
List with beta_hat, dfbeta, sigma_D, n, d, log_scale,
measure; or NULL if unsupported/degenerate/non-convergent.
Propensity adjustment is not represented here
This function has no ps_method or ps_adjust_method argument and reads
no sw / ps_hat / ips_covar column, so under ps_method != "none" it
fits the effect as if unadjusted rather than declining to fit it. It
cannot do otherwise: nothing in its signature tells it which case it is in.
The consequence is that the identifier ranks candidates on the IPTW or
g-computation adjusted effect built by make_effect_estimator(), while
every quantity derived here – \hat\beta(g), db_{g,i},
\sigma_{D,g}, and therefore the whole multiplier-resampling
correction – describes the unadjusted functional. forestsearch() also
forbids adjust_covariates alongside ps_method != "none", so the model
fitted here is fully unadjusted, not differently adjusted.
Earlier versions of this documentation claimed that propensity-adjusted
effects "return NULL (the caller falls back to splitting)". That was never
true and was not implementable as written. Whether to extend the influence
path, add the argument purely so the claim can be honoured, or refuse the
combination outright is unresolved.
Consistency proportion via literal splitting (single-stage)
Description
Runs n.splits Bernoulli splits through run_single_consistency_split()
and returns the rounded proportion of consistent splits, or NA_real_ when
too few valid splits are obtained. Survival and GLM (estimator_fn) paths.
Usage
.consistency_via_splits(
df.x,
N.x,
n.splits,
hr.consistency,
cox_init,
estimator_fn,
consistency_threshold,
adjust_covariates,
pconsistency.digits,
m = NA,
details = FALSE
)
Internal Workhorse for Single-Theta Detection Probability
Description
Internal Workhorse for Single-Theta Detection Probability
Usage
.detect_prob_glm_single(beta, sigma2_s, k1, k2, method, n_mc, seed)
Arguments
beta |
Numeric. True effect on natural parameter scale. |
sigma2_s |
Numeric. Split-half variance (8 / d_eff). |
k1 |
Numeric. Screening threshold (natural scale). |
k2 |
Numeric. Consistency threshold (natural scale). |
method |
Character. Integration method. |
n_mc |
Integer. Monte Carlo samples. |
seed |
Integer. RNG seed. |
Value
Numeric probability.
Fit Separate Regressions by Treatment Arm
Description
Fit Separate Regressions by Treatment Arm
Usage
.fit_arms(
Y,
W,
X,
poly_order,
covariate_select,
t_threshold,
regression = "ols",
offset = NULL
)
Apply multiplier resampling with forestsearch's defaulting rules.
Description
Thin wrapper around fs_mr_inference() shared by the consistency, DINA, and
GRF branches so the call (and mr_inference_args overrides) stays identical
across methods. Wrapped in tryCatch so a reconstruction failure yields
NULL (the branch's normal output is unaffected) rather than aborting.
Usage
.fs_apply_mr(
df,
candidates,
selected_members,
spec,
admission,
effect_neighborhood,
reselection_default,
selection_rule_default = "neighborhood",
mr_inference_args = list(),
seedit = NULL
)
Arguments
reselection_default |
Method-appropriate default re-selection rule
( |
Forwarded argument set
The wrapper forwards the full fs_mr_inference() argument set, so the
DINA and GRF branches are controllable through mr_inference_args exactly as
the consistency branch is through its own direct call. Before this, ten
arguments were dropped here – return_reselection, field_R_out,
field_R_in, field_uniform, field_M_cap, field_complement,
field_decompose, field_scale_complement, ij_residual, field_recovery
– which left field_decompose and field_recovery unreachable on
those engines and the other three inert while their defaults happened
to coincide with the intended values: a run's meta recorded them as set while
they controlled nothing.
Defaults are unchanged on every branch. Each newly forwarded
argument defaults to fs_mr_inference()'s own formal default, read from
formals(fs_mr_inference) at call time rather than restated here, so the two
cannot drift and a call that asks for nothing is byte-identical to the same
call before the change. match.arg()-style defaults are forwarded as their
whole candidate vector, which fs_mr_inference() resolves exactly as it would
have resolved a missing argument. ci_method keeps this wrapper's own
"ij" default, which is deliberately not the consistency branch's
"field"; changing it is a separate decision.
Assemble the influence matrix from candidate membership lists
Description
For each candidate (a vector of row indices into df), fit the treatment
effect once and place its dfbeta into the candidate's rows (zero elsewhere).
Candidates whose fit fails or whose dfbeta length does not match are dropped.
Usage
.fs_mr_assemble(df, candidates, spec)
Arguments
df |
Analysis data frame. |
candidates |
Named list of integer row-index vectors. |
spec |
See |
Value
List with B (N x S_kept), beta_hat, sigma_D, sizes,
names, keep, log_scale.
Default near-null harm-confirmation threshold for a measure
Description
The value t_confirm takes when left NULL: the null effect on the effect
scale, i.e. no effect at all rather than the screening threshold.
Usage
.fs_mr_confirm_null(effect_measure)
Arguments
effect_measure |
Character measure label ( |
Value
1 for ratio measures (effect scale), 0 for differences.
Build candidate-membership lists from a (v1,d1,c1,v2,d2,c2) table
Description
Build candidate-membership lists from a (v1,d1,c1,v2,d2,c2) table
Usage
.fs_mr_family_from_table(df, tab, op_right = ">=", n_min = 1L, digits = 17L)
Arguments
df |
Analysis data frame (for membership evaluation). |
tab |
Data frame with columns |
op_right |
Operator used when direction is not |
n_min |
Minimum subgroup size; smaller candidates are dropped. |
digits |
Threshold label formatting (matches the package convention). |
Value
Named list of integer row-index vectors (one per candidate).
Infinitesimal-jackknife variance of the de-biased (bagged) estimate
Description
Implements Leon et al. (2024) Eq. (VInfJ)/(VInfJ_bc) in multiplier form.
The multiplier matrix Xi supplies the centered bootstrap multiplicities
K^*_{bi} - \bar K^*_i, and r is the per-draw residual
r_b = \hat\beta(\widehat H) - \eta^*_b(\widehat H^*_b) - \eta^*_b(\widehat H) - \hat\beta^*(\widehat H),
i.e. (selection_bias + fixed_bias) - D_{H*_b}(b) - D_{H}(b), evaluated only
on the draws in ok that produced a re-selected winner.
Usage
.fs_mr_ij_var(Xi, r, ok)
Arguments
Xi |
|
r |
Length- |
ok |
Integer indices of usable draws. |
Value
List with tilde_V (raw IJ), hat_V (Wager 2014 bias-corrected),
and B_ok (number of usable draws).
Row-index membership for one candidate conjunction.
Description
Row-index membership for one candidate conjunction.
Usage
.fs_mr_members_from_conj(df, cj)
Arguments
df |
Analysis data frame. |
cj |
Conjunction data frame with columns |
Value
Integer row indices of subjects in the (harm) subgroup.
Mean-zero unit-variance multiplier vectors (one column per draw)
Description
Mean-zero unit-variance multiplier vectors (one column per draw)
Usage
.fs_mr_multipliers(n, draws, type)
Per-candidate influence pieces (Cox / GLM dispatch)
Description
Thin dispatch onto the existing forestsearch internals so MR uses the
identical treatment dfbeta, beta_hat, and robust sigma_D the resample
consistency engine uses.
Usage
.fs_mr_pieces(df_sub, spec)
Arguments
df_sub |
Data frame for one candidate subgroup's members. |
spec |
List with |
Value
Output of .consistency_glm_pieces() / .consistency_cox_pieces(),
or NULL.
Map sg_focus to MR's re-selection rule, faithfully per engine.
Description
MR must re-select under the same rule the search used. The only rule
that differs by engine is the "hr"/"eff" focus: the consistency search
ranks by consistency rate, whose MR analog is "maxcons"; DINA/GRF rank by
effect, whose analog is "maxeff". Every other focus maps identically
(maxSG/minSG pass through; hrMaxSG/hrMinSG -> the size-within-effect-
neighborhood rules effMaxSG/effMinSG). This makes sg_focus the single
source of truth, so callers never need to specify reselection by hand.
Usage
.fs_mr_reselection_from_focus(sg_focus, engine = c("effect", "consistency"))
Arguments
sg_focus |
The search focus ( |
engine |
|
Resolve a de-biased SE from the IJ variance, with graceful fallback
Description
Prefers the bias-corrected IJ variance; if it is non-positive (too few draws), falls back to the raw IJ variance, then to the subgroup robust SE.
Usage
.fs_mr_se_from_ij(ij, se_fallback)
Arguments
ij |
Output of |
se_fallback |
Robust subgroup SE ( |
Value
List with se, var, and source
("ij", "ij_raw", or "wald_fallback").
Apply a selection rule among passing candidates on one draw
Description
For the size-within-band rules (effMaxSG/effMinSG) the inclusion band is
built to match the search's selection_rule:
-
"neighborhood"- effect withinnbhdof the max, on the NATURAL effect scale and multiplicative (eff >= (1 - nbhd) * max(eff)), matchingsort_subgroups()exactly (G3); -
"pareto"- the 2-D non-dominated set in (effect, size), via the same dominance core the search uses (.pareto_dominated_xy()) (G2); -
"both"- the intersection of the two.betais on the working scale (log for ratio measures);log_scalecontrols the conversion to the natural effect for the band.
Usage
.fs_mr_select(
beta,
zcons,
sizes,
passers,
rule,
nbhd,
selection_rule = "neighborhood",
log_scale = TRUE
)
Build the Cox/GLM spec for MR from forestsearch()'s resolved columns.
Description
DINA/GRF operate on the user-facing analysis frame df (not the standardized
df.fs), so survival uses the user column names here – unlike the
consistency hook, which uses the internal Y/Event/Treat names on df.fs.
Usage
.fs_mr_spec(
outcome_type,
effect_measure,
treat.name,
outcome.name,
event.name,
offset.name,
adjust_covariates,
adverse_outcome,
df
)
Can the GLM resampling approximation represent this effect measure?
Description
The resampling approximation requires the treatment effect to be a single model coefficient with a well-defined influence function. That holds for OR/RR (logistic / log-binomial), RD (identity-link binomial), MD (OLS), and IRR (Poisson with offset). It does not hold for IRD (a delta-method rate difference) or propensity-adjusted (IPTW / G-computation) effects, for which the caller falls back to literal splitting.
Usage
.glm_resample_supported(spec)
Arguments
spec |
List; the |
Value
Logical scalar.
Re-format Numeric Cells in FSsg_tab at a Lower Precision
Description
Parses formatted-string cells in FSsg_tab back to numerics and
re-formats at the requested digits. Designed for display-time
precision control in summarize_bootstrap_results.
Usage
.reformat_FSsg_tab_digits(tab, digits)
Arguments
tab |
Data frame or matrix with formatted-string cells. |
digits |
Integer. Decimal places. |
Details
Cells are identified by column name pattern, not content sniffing. Three cell formats are recognized:
Single value with optional sign and percent suffix (Mean(C)/(T), Rate(C)/(T), Diff):
"-56.4321","+12.5","42.1%".Three-value CI (effect-measure column,
"\{measure\} (95% CI)"):"82.7831 (14.8067, 150.7585)".Three-value CI without space (bias-corrected
"\{measure\}*"column):"72.4500 (-61.5400,206.4400)".
Columns NOT touched: Subgroup (text label), n and
n1 (count-with-percent format like "87 (8.0%)" —
different semantics).
Robustness contract: cells that don't match any expected pattern are
returned unchanged. Construction-time precision must be at least
digits for the reformat to be informative; if construction was
at lower precision, this function will zero-pad to digits but
cannot recover information.
Value
Reformatted tab, same shape.
Select Covariates for the Crump Test
Description
Returns a list with X (selected columns) and selected_names.
Usage
.select_covariates(Y, W, X, method, threshold)
Cross-Validation Subgroup Match Summary
Description
Summarizes the match between cross-validation subgroups and analysis subgroups.
Usage
CV_sgs(sg1, sg2, confs, sg_analysis)
Arguments
sg1 |
Character vector. Subgroup 1 labels for each fold. |
sg2 |
Character vector. Subgroup 2 labels for each fold. |
confs |
Character vector. Confounder names. |
sg_analysis |
Character vector. Subgroup analysis labels. |
Value
List with indicators for any match, exact match, one match, and covariate-specific matches.
Examples
## Not run:
# CV_sgs() is called internally by forestsearch_KfoldOut().
# See forestsearch_Kfold() and forestsearch_KfoldOut() for the entry points.
## End(Not run)
Convert Factor Code to Label
Description
Converts q-indexed codes to human-readable labels using the confs_labels mapping. Supports both full format ("q1.1", "q3.0") and short format ("q1", "q3"). Handles vector input via recursion.
Usage
FS_labels(Qsg, confs_labels)
Arguments
Qsg |
Character. Factor code in format |
confs_labels |
Character vector. Labels for each factor, indexed by factor number. |
Value
Character. Human-readable label wrapped in braces, e.g.,
"{age <= 50}" or "!{age <= 50}" for complement. Returns the
original code if no match is found.
Examples
## Not run:
labels <- c("age <= 50", "tumor size <= 20", "nodes <= 3")
FS_labels("q1.1", labels)
FS_labels("q1.0", labels)
FS_labels("q2", labels)
FS_labels(c("q1", "q3"), labels)
## End(Not run)
Subgroup summary table estimates
Description
Returns a summary table of subgroup estimates (HR, RMST, medians, etc.).
Usage
SG_tab_estimates(
df,
SG_flag,
outcome.name = "tte",
event.name = "event",
treat.name = "treat",
strata.name = NULL,
hr_1a = NA,
hr_0a = NA,
potentialOutcome.name = NULL,
sg1_name = NULL,
sg0_name = NULL,
draws = 0,
details = FALSE,
return_medians = TRUE,
est.scale = "hr"
)
Arguments
df |
Data frame. |
SG_flag |
Character. Subgroup flag variable. |
outcome.name |
Character. Name of outcome variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
strata.name |
Character. Name of strata variable (optional). |
hr_1a |
Character. Adjusted HR for subgroup 1 (optional). |
hr_0a |
Character. Adjusted HR for subgroup 0 (optional). |
potentialOutcome.name |
Character. Name of potential outcome variable (optional). |
sg1_name |
Character. Name for subgroup 1. |
sg0_name |
Character. Name for subgroup 0. |
draws |
Integer. Number of draws for resampling (optional). |
details |
Logical. Print details. |
return_medians |
Logical. Use medians or RMST. |
est.scale |
Character. Effect scale ("hr" or "1/hr"). |
Value
Data frame of subgroup summary estimates.
Examples
## Not run:
library(survival)
df <- data.frame(
tte = gbsg$rfstime / 30.4375,
event = gbsg$status,
treat = gbsg$hormon,
treat.recommend = as.integer(gbsg$er > 0)
)
SG_tab_estimates(df, SG_flag = "ITT",
outcome.name = "tte",
event.name = "event",
treat.name = "treat")
## End(Not run)
Subgroup Summary Table for GLM Outcomes
Description
GLM counterpart to SG_tab_estimates. Produces a summary data frame
with per-arm outcome rates and the GLM effect estimate for each subgroup.
Usage
SG_tab_estimates_glm(
df,
SG_flag,
outcome.name,
treat.name,
estimator_fn = NULL,
effect_measure = "RD",
outcome_type = "binary",
effect_a_1 = NA,
effect_a_0 = NA,
sg1_name = "Recommend",
sg0_name = "Questionable",
est.scale = "hr",
digits = 4
)
Arguments
df |
Data frame with a |
SG_flag |
Character. "ITT" for the full population, or the name of the subgroup flag column (e.g., "treat.recommend"). |
outcome.name |
Character. Name of outcome variable. |
treat.name |
Character. Name of treatment variable. |
estimator_fn |
Closure from |
effect_measure |
Character. Effect measure label (e.g., "RD", "OR"). |
outcome_type |
Character. One of |
effect_a_1 |
Character. Adjusted effect for subgroup 1 (optional). |
effect_a_0 |
Character. Adjusted effect for subgroup 0 (optional). |
sg1_name |
Character. Label for subgroup 1 (treat.recommend == 1). |
sg0_name |
Character. Label for subgroup 0 (treat.recommend == 0). |
est.scale |
Character. Effect scale ("hr" or "1/hr"). |
digits |
Integer. Decimal places for numeric formatting; passed
through to |
Value
Data frame. One row when SG_flag = "ITT"; two rows
(sg0 then sg1) when SG_flag = "treat.recommend". Columns:
Subgroup, n, n1, Rate(C) or
Mean(C) (binary vs continuous/count), Rate(T) or
Mean(T), Diff, and the effect-measure header
"\{effect_measure\} (95% CI)". When effect_a_* is
supplied, an additional "\{effect_measure\}*" column carries
the bias-corrected estimate.
Violin/Boxplot Visualization of HR Estimates
Description
Creates violin plots with embedded boxplots showing the distribution of hazard ratio estimates across simulations for different analysis populations. Supports symmetric trimming to handle extreme values that can distort the display when small subgroups produce very large HR estimates.
Usage
SGplot_estimates(
df,
label_training = "Training",
label_testing = "Testing",
label_itt = "ITT (stratified)",
label_sg = "Testing (subgroup)",
trim_fraction = NULL,
ylim = NULL,
show_summary = NULL,
title = "Distribution of HR Estimates Across Simulations",
subtitle = NULL
)
Arguments
df |
data.frame or data.table. Simulation results from
|
label_training |
Character. Label for training data estimates. Default: "Training" |
label_testing |
Character. Label for testing data estimates. Default: "Testing" |
label_itt |
Character. Label for ITT estimates. Default: "ITT (stratified)" |
label_sg |
Character. Label for subgroup estimates. Default: "Testing (subgroup)" |
trim_fraction |
Numeric or NULL. Fraction of observations to trim from each tail (e.g., 0.01 trims the lowest 1\ When non-NULL, trimmed means and SDs are computed for each group, extreme observations are flagged, and the y-axis is clipped to the trimmed data range. Set to NULL (default) for no trimming (backward compatible). |
ylim |
Numeric vector of length 2 or NULL. Explicit y-axis limits
as |
show_summary |
Logical. Annotate each violin with mean (SD) below
the x-axis labels. When trimming is active, displays trimmed
statistics. Default: TRUE when |
title |
Character. Plot title. Default: "Distribution of HR Estimates Across Simulations". |
subtitle |
Character or NULL. Plot subtitle. When trimming is active and subtitle is NULL, an auto-generated note indicating the trim fraction and number of flagged observations is shown. Default: NULL. |
Value
List with components:
- dfPlot_estimates
data.table formatted for plotting, with a
trimmedlogical column when trimming is active- plot_estimates
ggplot2 object
- trim_info
List of per-group trimming diagnostics (NULL when no trimming). Each element contains:
n_total,n_trimmed,n_flagged,raw_mean,raw_sd,trimmed_mean,trimmed_sd,lower_bound,upper_bound.
See Also
mrct_region_sims for generating simulation results,
summaryout_mrct for tabular summaries with trimming
Examples
## Not run:
# Default (no trimming) — backward compatible
plot_results <- SGplot_estimates(
results_alt,
label_training = "Non-Region A, ITT",
label_itt = "Overall, ITT",
label_testing = "Region A, ITT",
label_sg = "Region A, identified subgroup"
)
print(plot_results$plot_estimates)
# With 1% symmetric trimming
plot_results <- SGplot_estimates(
results_alt,
label_training = "Non-AP, ITT",
label_itt = "Overall, ITT",
label_testing = "AP, ITT",
label_sg = "AP, identified subgroup",
trim_fraction = 0.01
)
print(plot_results$plot_estimates)
# Inspect trimming diagnostics
print(plot_results$trim_info)
# With explicit y-axis limits
plot_results <- SGplot_estimates(
results_alt,
trim_fraction = 0.01,
ylim = c(0.2, 3.0)
)
## End(Not run)
Disjunctive (dummy) coding for factor columns
Description
Disjunctive (dummy) coding for factor columns
Usage
acm.disjctif(df)
Arguments
df |
Data frame with factor variables. |
Value
Data frame with dummy-coded columns.
Add ID Column to Data Frame
Description
Ensures that a data frame has a unique ID column. If id.name is not
provided, a column named "id" is added. If id.name is provided
but does not exist in the data frame, it is created with unique integer
values.
Usage
add_id_column(df.analysis, id.name = NULL)
Arguments
df.analysis |
Data frame to which the ID column will be added. |
id.name |
Character. Name of the ID column to add (default is
|
Value
Data frame with the ID column added if necessary.
Add Unprocessed Variables from Original Data
Description
Add Unprocessed Variables from Original Data
Usage
add_unprocessed_vars(
df_work,
data,
outcome_var,
event_var,
treatment_var,
continuous_vars,
factor_vars,
verbose
)
Covariate set handed to forestsearch()/mrct_region_sims() for the paper's observed vs unobserved analysis cases.
Description
Covariate set handed to forestsearch()/mrct_region_sims() for the paper's observed vs unobserved analysis cases.
Usage
analysis_covariates(
dgm,
case = c("observed", "unobserved"),
drop = character(0)
)
Arguments
dgm |
An MRCT data-generating mechanism returned by
|
case |
Character. |
drop |
Character vector of covariate names (without the |
Value
Character vector of z_-prefixed covariate column names.
Analyze subgroup for summary table (OPTIMIZED)
Description
Analyzes a subgroup and returns formatted results for summary table. Uses optimized cox_summary() and reduces redundant calculations.
Usage
analyze_subgroup(
df_sub,
outcome.name,
event.name,
treat.name,
strata.name,
subgroup_name,
hr_a,
potentialOutcome.name,
return_medians,
N
)
Arguments
df_sub |
Data frame for subgroup. |
outcome.name |
Character. Name of outcome variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
strata.name |
Character. Name of strata variable (optional). |
subgroup_name |
Character. Subgroup name. |
hr_a |
Character. Adjusted hazard ratio (optional). |
potentialOutcome.name |
Character. Name of potential outcome variable (optional). |
return_medians |
Logical. Use medians or RMST. |
N |
Integer. Total sample size. |
Value
Character vector of results.
Examples
## Not run:
library(survival)
df <- data.frame(
tte = gbsg$rfstime / 30.4375,
event = gbsg$status,
treat = gbsg$hormon
)
analyze_subgroup(
df_sub = df,
outcome.name = "tte",
event.name = "event",
treat.name = "treat",
strata.name = NULL,
subgroup_name = "All",
hr_a = NA,
potentialOutcome.name = NULL,
return_medians = TRUE,
N = nrow(df)
)
## End(Not run)
Analyze subgroup for GLM outcomes
Description
GLM counterpart to analyze_subgroup. Computes per-arm outcome
rates, sample sizes, and the treatment effect estimate (RD, OR, etc.)
using the estimator closure.
Usage
analyze_subgroup_glm(
df_sub,
outcome.name,
treat.name,
subgroup_name,
effect_a = NA,
estimator_fn = NULL,
effect_measure = "RD",
outcome_type = "binary",
N,
digits = 4
)
Arguments
df_sub |
Data frame for the subgroup. |
outcome.name |
Character. Name of outcome variable. |
treat.name |
Character. Name of treatment variable. |
subgroup_name |
Character. Label for this subgroup. |
effect_a |
Character. Adjusted effect estimate string (optional). |
estimator_fn |
Closure from |
effect_measure |
Character. Effect measure label (e.g., "RD", "OR"). |
outcome_type |
Character. One of |
N |
Integer. Total sample size (for percentage calculation). |
digits |
Integer. Number of decimal places for arm-mean / Diff /
effect-estimate / CI numeric formatting. Default 4 to provide
headroom for downstream display reformatting (e.g., via
|
Value
One-row data frame of formatted subgroup results. Column names
are placeholders (Subgroup, n, n1, rate0,
rate1, Diff, effect, optionally effect_a)
and are overwritten by the calling function
(SG_tab_estimates_glm) with the user-facing labels
(e.g., "Rate(C)" vs "Mean(C)", the effect-measure
header). Returning a data frame – rather than a named character
vector via c(...) – ensures type stability across the ITT
and subgroup call paths and prevents downstream rendering failures
in consumers that go through as.data.frame() (e.g., gt).
Apply Spline Constraint to Treatment Effect Coefficients
Description
Apply Spline Constraint to Treatment Effect Coefficients
Usage
apply_spline_constraint(b0, spline_var, knot, zeta, log_hrs, k_treat, verbose)
Assemble Final Results Object
Description
Assemble Final Results Object
Usage
assemble_results(
df_super,
mu,
tau,
gamma,
b0,
cens_model,
subgroup_vars,
subgroup_cuts,
subgroup_definitions,
hr_results,
continuous_vars,
factor_vars,
model,
n_super,
seed,
spline_info = NULL,
df_source = NULL
)
Assign data to subgroups based on selected node
Description
Creates treatment recommendation flags based on identified subgroup
Usage
assign_subgroup_membership(data, best_subgroup, trees, X)
Arguments
data |
Data frame. Original data |
best_subgroup |
Data frame row. Selected subgroup information |
trees |
List. Policy trees |
X |
Matrix. Covariate matrix |
Value
Data frame with added predict.node and treat.recommend columns
Random-benchmark subgroup specification
Description
Describes the size-matched random benchmark subgroups regenerated in
every simulated trial. With the default nested = TRUE, a single
index draw of max(sizes) patients is taken per trial and the smaller
benchmarks are its prefixes (random15 \subset random20
\subset random40 \subset random60), exactly as in the
extreme-subgroups vignettes. Membership columns are named
paste0(prefix, sizes) and take values 0/1.
Usage
benchmark_spec(sizes = c(60L, 40L, 20L, 15L), nested = TRUE, prefix = "random")
Arguments
sizes |
Integer vector of benchmark sizes. Order determines the
column-creation order only; the draw uses |
nested |
Logical. |
prefix |
Column-name prefix, default |
Value
An object of class "benchmark_spec".
See Also
Examples
benchmark_spec()
benchmark_spec(sizes = c(100, 50), prefix = "rnd")
Bootstrap Results for ForestSearch with Bias Correction
Description
Runs bootstrap analysis for ForestSearch, fitting Cox models and computing bias-corrected estimates and valid CIs (see vignette for references)
Usage
bootstrap_results(
fs.est,
df_boot_analysis,
cox.formula.boot,
nb_boots,
show_three,
H_obs,
Hc_obs,
seed = 8316951L,
estimator_fn = NULL,
effect_measure = NULL,
boot_index_mat = NULL,
mr_in_replicates = FALSE
)
Arguments
fs.est |
List. ForestSearch results object from
|
df_boot_analysis |
Data frame. Bootstrap analysis data with same structure
as |
cox.formula.boot |
Formula. Cox model formula for bootstrap, typically
created by |
nb_boots |
Integer. Number of bootstrap samples to generate (e.g., 500-1000). More iterations provide better bias correction but increase computation time. |
show_three |
Logical. If |
H_obs |
Numeric. Observed log hazard ratio for subgroup H (harm/questionable group,
|
Hc_obs |
Numeric. Observed log hazard ratio for subgroup H^c (complement/recommend,
|
seed |
Integer. Random seed for reproducibility. Default 8316951L.
Only used when |
estimator_fn |
Closure or |
effect_measure |
Character or |
boot_index_mat |
Integer matrix or |
mr_in_replicates |
Logical. Whether multiplier resampling runs inside
each replicate. Default |
Value
Data.table with one row per bootstrap iteration and columns:
- boot_id
Integer. Bootstrap iteration number (1 to
nb_boots)- H_biasadj_1
Bias-corrected estimate for H using method 1:
H_obs - (Hstar_star - Hstar_obs)- H_biasadj_2
Bias-corrected estimate for H using method 2:
2*H_obs - (H_star + Hstar_star - Hstar_obs)- Hc_biasadj_1
Bias-corrected estimate for H^c using method 1
- Hc_biasadj_2
Bias-corrected estimate for H^c using method 2
- max_sg_est
Numeric. Maximum subgroup hazard ratio found
- L
Integer. Number of candidate factors evaluated
- max_count
Integer. Maximum number of factor combinations
- events_H_0
Integer. Number of events in control arm of original subgroup H on bootstrap sample
- events_H_1
Integer. Number of events in treatment arm of original subgroup H on bootstrap sample
- events_Hc_0
Integer. Number of events in control arm of original subgroup H^c on bootstrap sample
- events_Hc_1
Integer. Number of events in treatment arm of original subgroup H^c on bootstrap sample
- events_Hstar_0
Integer. Number of events in control arm of new subgroup H* on original data
- events_Hstar_1
Integer. Number of events in treatment arm of new subgroup H* on original data
- events_Hcstar_0
Integer. Number of events in control arm of new subgroup H^c* on original data
- events_Hcstar_1
Integer. Number of events in treatment arm of new subgroup H^c* on original data
- tmins_search
Numeric. Minutes spent on subgroup search in this iteration
- tmins_iteration
Numeric. Total minutes for this bootstrap iteration
- Pcons
Numeric. Consistency p-value for top subgroup
- any_found
Integer 0/1. Identification flag; equivalent to
as.integer(!is.na(Pcons)). Added in v0.2.0 to mirror thefold_summary$any_foundcolumn produced byforestsearch_tenfold, so cross-API diagnostics (sum(results$any_found == 1L)) work identically for bootstrap and CV outputs.- hr_sg
Numeric. Hazard ratio for top subgroup
- N_sg
Integer. Sample size of top subgroup
- E_sg
Integer. Number of events in top subgroup
- K_sg
Integer. Number of factors defining top subgroup
- g_sg
Numeric. Subgroup group ID
- m_sg
Numeric. Subgroup index
- M.1
Character. First factor label
- M.2
Character. Second factor label
- M.3
Character. Third factor label
- M.4
Character. Fourth factor label
- M.5
Character. Fifth factor label
- M.6
Character. Sixth factor label
- M.7
Character. Seventh factor label
- grf_cuts_b
Character. GRF policy-tree cut expressions returned on this bootstrap sample, collapsed with " | " when multiple cuts are produced.
NA_character_when GRF was not used, the bootstrap forestsearch call errored, or GRF returned no cuts. Populated independently of whether a subgroup was identified: GRF may surface a candidate cut on bootstraps where the downstream consistency stage rejects all candidates. Mirrors thegrf_cutscolumn thatforestsearch_tenfoldcaptures in itsfold_summaryreturn slot.
Rows where no valid subgroup was found will have NA for bias corrections.
The returned object has a "timing" attribute with summary statistics.
Bias Correction Methods
Two bias correction approaches are implemented:
-
Method 1 (Simple Optimism):
H_{adj1} = H_{obs} - (H^*_{*} - H^*_{obs})where
H^*_{*}is the new subgroup HR on bootstrap data andH^*_{obs}is the new subgroup HR on original data. -
Method 2 (Double Bootstrap):
H_{adj2} = 2 \times H_{obs} - (H_{*} + H^*_{*} - H^*_{obs})where
H_{*}is the original subgroup HR on bootstrap data.
where:
-
H_obs: Original subgroup HR on original data -
H_star: Original subgroup HR on bootstrap data -
Hstar_obs: New subgroup (found in bootstrap) HR on original data -
Hstar_star: New subgroup (found in bootstrap) HR on bootstrap data
Computational Details
Uses
doFuturebackend for parallel execution (configured externally)Sets reproducible seeds:
8316951 + boot * 100for each iterationEach bootstrap iteration runs full ForestSearch pipeline including variable selection, subgroup search, and consistency evaluation
Sequential execution within each bootstrap prevents nested parallelization
Failed bootstrap iterations generate warnings but don't stop execution
Confounders are removed from bootstrap data to force fresh variable selection
Bootstrap Configuration
Each bootstrap iteration modifies ForestSearch arguments to:
-
Suppress output:
details,show_candidate_summary,plot.sg,plot.grfall set toFALSE -
Force re-selection:
grf_resandgrf_cutsset toNULL -
Prevent nested parallel:
parallel_args$plan = "sequential",workers = 1
Performance Considerations
Typical runtime: 1-5 seconds per bootstrap iteration
For 1000 bootstraps with 6 workers: ~3-10 minutes total
Memory usage scales with dataset size and number of workers
Consider reducing
nb_bootsfor initial testing (e.g., 100)
Error Handling
The function gracefully handles three failure modes:
Bootstrap sample creation fails: Returns row with all
NAForestSearch fails to run: Warns and returns row with all
NAForestSearch runs but finds no subgroup: Returns row with all
NA
All three cases ensure the foreach loop can still combine results via rbind.
Note
This function is designed to be called within a foreach loop
with %dofuture% operator. It requires:
All functions in
get_bootstrap_exportsto be available in the parallel workersPackages listed in
BOOTSTRAP_REQUIRED_PACKAGESto be installedProper parallel backend setup via
setup_parallel_SGcons
See Also
forestsearch_bootstrap_dofuture for the wrapper function that
sets up parallelization and calls this function
build_cox_formula for creating the Cox formula
fit_cox_models for initial Cox model fitting
get_Cox_sg for Cox model fitting on subgroups
get_dfRes for processing bootstrap results into confidence intervals
bootstrap_ystar for generating the Ystar matrix
Examples
## Not run:
# Typically called via forestsearch_bootstrap_dofuture()
# Manual usage for debugging:
# 1. Fit initial ForestSearch model
fs_result <- forestsearch(
df.analysis = mydata,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
confounders.name = c("age", "sex", "stage")
)
# 2. Build Cox formula
cox_formula <- build_cox_formula("time", "status", "treatment")
# 3. Get observed estimates
cox_fits <- fit_cox_models(fs_result$df.est, cox_formula)
# 4. Set up parallel backend
library(doFuture)
plan(multisession, workers = 6)
# 5. Run bootstrap (note: this is already parallelized internally)
boot_results <- bootstrap_results(
fs.est = fs_result,
df_boot_analysis = fs_result$df.est,
cox.formula.boot = cox_formula,
nb_boots = 100,
show_three = TRUE,
H_obs = cox_fits$H_obs,
Hc_obs = cox_fits$Hc_obs
)
# 6. Check results
summary(boot_results)
# Proportion of bootstraps that found a subgroup
mean(!is.na(boot_results$H_biasadj_2))
## End(Not run)
Bootstrap Ystar Count Matrix
Description
Returns the B \times N integer count matrix consumed by the
infinitesimal-jackknife variance estimator in get_dfRes.
Entry (b, j) is the number of times row j of df was
selected on bootstrap iteration b.
Usage
bootstrap_ystar(df, nb_boots, seed = 8316951L)
Arguments
df |
Data frame. Counts are computed by positional row index;
|
nb_boots |
Integer. Number of bootstrap samples. |
seed |
Integer or |
Value
Integer matrix with nb_boots rows and nrow(df)
columns.
Implementation note (v0.2.0)
This function previously generated the count matrix via a parallel
foreach pass, with counts derived by matching against a
user-supplied id column. It now runs on the main process and
counts by positional row index via tabulate,
matching how bootstrap_results constructs bootstrap data
(df[in_boot, ]). For data with unique subject ids (the standard
clinical-trial case) the two implementations are equivalent. The new
implementation is also robust to non-unique id values.
Reproducibility note
Bootstrap indices generated by this version are not bit-identical to those produced by v0.1.x for the same seed (the generator now uses the main-process Mersenne-Twister stream rather than per-task L'Ecuyer streams). Statistical properties of bootstrap summaries are preserved.
Examples
## Not run:
df <- data.frame(id = 1:50, tte = rexp(50), event = rbinom(50, 1, 0.6),
treat = rep(0:1, 25))
ystar <- bootstrap_ystar(df, nb_boots = 10)
dim(ystar) # 10 x 50
rowSums(ystar) # all equal to 50
## End(Not run)
Build Polynomial Basis Matrix
Description
Build Polynomial Basis Matrix
Usage
build_basis(X, poly_order = 1L, include_intercept = TRUE)
Arguments
X |
Numeric matrix (N x d). Must be numeric (no factors). |
poly_order |
Integer. 1 = linear, 2 = squares + interactions. |
include_intercept |
Logical. Default: TRUE. |
Value
Numeric matrix (N x K).
Build Classification Rate Table from Simulation Results
Description
Constructs a publication-quality gt table summarizing subgroup
identification and classification rates across one or more data generation
scenarios and analysis methods. The layout mirrors Table 4 of
Leon et al. (2024) with metrics grouped by model scenario (null / alt)
and columns for each analysis method.
Usage
build_classification_table(
scenario_results,
analyses = NULL,
digits = 2,
title = "Subgroup Identification and Classification Rates",
n_sims = NULL,
bold_threshold = 0.05,
font_size = 12,
subgroup_notation = c("harm", "benefit")
)
Arguments
scenario_results |
Named list. Each element is itself a list with:
|
analyses |
Character vector of analysis labels to include
(e.g., |
digits |
Integer. Decimal places for proportions. Default: 2. |
title |
Character. Table title. Default:
|
n_sims |
Integer. Number of simulations (for subtitle). Default:
|
bold_threshold |
Numeric. Type I error threshold above which the
|
font_size |
Numeric. Font size in pixels for table text. Default: 12. Increase to 14 or 16 for larger display. |
subgroup_notation |
Character. |
Details
For each scenario the function computes:
-
any(H): Proportion of simulations identifying any subgroup. -
sens(H): Mean sensitivity (only under alternative). -
sens(Hc): Mean specificity. -
ppv(H): Mean positive predictive value (only under alternative). -
ppv(Hc): Mean negative predictive value. -
avg|H|: Mean size of identified subgroup (when found).
Under the null hypothesis the rows are reduced to any(H),
sens(Hc), ppv(Hc), and avg|H|.
Value
A gt table object.
See Also
format_oc_results,
summarize_simulation_results
Examples
## Not run:
# Assemble results from H0 and H1 simulations
scenarios <- list(
null = list(
results = results_null, label = "M1",
n_sample = 700, dgm = dgm_null, hypothesis = "null"
),
alt = list(
results = results_alt, label = "M1",
n_sample = 700, dgm = dgm_calibrated, hypothesis = "alt"
)
)
build_classification_table(scenarios, n_sims = 100)
## End(Not run)
Build Cox Model Formula
Description
Constructs a Cox model formula from variable names, optionally adjusted for additional covariates.
Usage
build_cox_formula(
outcome.name,
event.name,
treat.name,
adjust_covariates = NULL
)
Arguments
outcome.name |
Character. Name of outcome variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
adjust_covariates |
Character vector or |
Value
An R formula object for Cox regression. When adjust_covariates
is supplied, the treatment term is always placed first on the
right-hand side.
Examples
build_cox_formula("time_months", "event", "treat")
build_cox_formula("time_months", "event", "treat",
adjust_covariates = c("strata(site)", "age"))
Build Estimation Properties Table from Simulation Results
Description
Constructs a publication-quality gt table summarizing estimation
properties for hazard ratios in the identified subgroup and its complement.
The layout mirrors Table 5 of Leon et al. (2024), showing average estimate,
empirical SD, min, max, and relative bias for each estimator.
Usage
build_estimation_table(
results,
dgm,
analysis_method = "FSlg",
n_boots = NULL,
digits = 2,
title = "Estimation Properties",
subtitle = NULL,
font_size = 12,
cde_H = NULL,
cde_Hc = NULL,
subgroup_notation = c("harm", "benefit"),
trim_threshold = 1000,
trim_fraction = 0.01
)
Arguments
results |
|
dgm |
DGM object. Used for true parameter values ( |
analysis_method |
Character. Which analysis method to tabulate
(e.g., |
n_boots |
Integer or |
digits |
Integer. Decimal places. Default: 2. |
title |
Character. Table title. |
subtitle |
Character or |
font_size |
Numeric. Font size in pixels for table text. Default: 12. Increase to 14 or 16 for larger display. |
cde_H |
Numeric or |
cde_Hc |
Numeric or |
subgroup_notation |
Character. |
trim_threshold |
Numeric or |
trim_fraction |
Numeric in |
Details
Uses the paper's notation conventions:
theta-dagger: Marginal (causal) HR truth
theta-ddagger: Controlled direct effect (CDE) truth
theta-hat(H-hat): Plugin Cox estimate in identified subgroup
theta-hat*(H-hat): Bootstrap bias-corrected estimate
Includes both Cox-based HR and AHR (Average Hazard Ratio from loghr_po) estimators when AHR columns are present in the results.
For each subgroup (H and Hc) the function reports:
-
Avg: Mean of the estimates across estimable simulations.
-
SD: Empirical standard deviation.
-
Min / Max: Range.
-
b-dagger: Relative bias (percent) vs marginal truth,
100 * (Avg - theta_dagger) / theta_dagger. -
b-ddagger (conditional): Relative bias (percent) vs CDE truth, shown when CDE values are available.
When bootstrap-corrected columns (hr.H.bc, hr.Hc.bc) are
present in results, an additional bias-corrected row
(theta-hat*(H-hat)) is added per subgroup.
When AHR columns (ahr.H.hat, ahr.Hc.hat) are present, AHR
estimation rows are appended using the DGM's true AHR values for relative
bias calculation.
When CDE columns (cde.H.hat, cde.Hc.hat) are present and
CDE truth values are available, CDE estimation rows
(theta-ddagger(H-hat)) are appended. The b-dagger column for CDE rows
reports bias relative to the CDE truth rather than the marginal HR.
Value
A gt table object, or NULL if no estimable
realizations exist.
See Also
build_classification_table,
format_oc_results, get_dgm_hr
Examples
## Not run:
# Basic usage (marginal truth only)
build_estimation_table(
results = results_alt,
dgm = dgm_calibrated,
analysis_method = "FSlg"
)
# With CDE truth (full Table 5 alignment)
build_estimation_table(
results = results_alt,
dgm = dgm_calibrated,
cde_H = 2.25,
cde_Hc = 0.60
)
## End(Not run)
Calculate Covariance for Bootstrap Estimates
Description
Calculates the covariance between a vector and bootstrap estimates.
Usage
calc_cov(x, Est)
Arguments
x |
Numeric vector. |
Est |
Numeric vector of bootstrap estimates. |
Value
Numeric value of covariance.
Calculate counts for subgroup summary
Description
Calculates sample size, treated count, and event count for a subgroup.
Usage
calculate_counts(Y, E, Treat, N)
Arguments
Y |
Numeric vector of outcome. |
E |
Numeric vector of event indicators. |
Treat |
Numeric vector of treatment indicators. |
N |
Integer. Total sample size. |
Value
List with formatted counts.
Calculate Event Counts by Treatment Arm
Description
Calculate Event Counts by Treatment Arm
Usage
calculate_event_counts(dd, tt, id.x)
Calculate Hazard Ratios from Potential Outcomes
Description
Calculate Hazard Ratios from Potential Outcomes
Usage
calculate_hazard_ratios(df_super, n_super, mu, tau, model, verbose)
Arguments
df_super |
Data frame with super population |
n_super |
Size of super population |
mu |
Intercept parameter |
tau |
Scale parameter |
model |
Model type ("alt" or "null") |
verbose |
Logical for verbose output |
Value
List of hazard ratios
Calculate Linear Predictors for Potential Outcomes
Description
Calculate Linear Predictors for Potential Outcomes
Usage
calculate_linear_predictors(
df_super,
covariate_cols,
gamma,
b0,
spline_info = NULL
)
Calculate Maximum Combinations
Description
Calculate Maximum Combinations
Usage
calculate_max_combinations(L, maxk)
Calculate potential outcome hazard ratio
Description
Calculates the average hazard ratio from a potential outcome variable.
Usage
calculate_potential_hr(df, potentialOutcome.name)
Arguments
df |
Data frame. |
potentialOutcome.name |
Character. Name of potential outcome variable. |
Value
Numeric value of average hazard ratio.
Calculate Skewness
Description
Helper function to calculate sample skewness.
Usage
calculate_skewness(x)
Arguments
x |
Numeric vector |
Value
Numeric skewness value
Calibrate L_eff from Simulation Results
Description
Given per-subgroup approximation probabilities and simulated FPR
values at multiple sample sizes, fits the power-law model
L_{\text{eff}}(N) = C \cdot (N / n_{\min})^\alpha.
Usage
calibrate_L_eff(N, P1, sim_fpr, n_min = 60)
Arguments
N |
Integer vector. Sample sizes at which simulations were run. |
P1 |
Numeric vector. Per-subgroup detection probability from
|
sim_fpr |
Numeric vector. Simulated procedure-level FPR at each N. |
n_min |
Integer. Minimum subgroup size used in ForestSearch. Default: 60. |
Details
At each sample size, the implied L_{\text{eff}} is:
L_{\text{eff}} = \frac{\log(1 - \text{FPR}_{\text{sim}})}
{\log(1 - P_1)}
A log-linear model is then fit:
\log(L_{\text{eff}}) = \log(C) + \alpha \log(N / n_{\min}).
Value
A list of class "leff_calibration" with components:
- C
Numeric. Fitted constant.
- alpha
Numeric. Fitted power-law exponent.
- n_min
Integer. Minimum subgroup size.
- r_squared
Numeric. R-squared of the log-linear fit.
- data
Data frame with N, P1, sim_fpr, L_eff, and fitted values.
- fit
The
lmobject.
Examples
# From binary threshold calibration results
cal <- calibrate_L_eff(
N = c(200, 500, 700, 1000),
P1 = c(0.152, 0.109, 0.091, 0.071),
sim_fpr = c(0.185, 0.220, 0.270, 0.280),
n_min = 60
)
print(cal)
Calibrate Censoring Adjustment to Match DGM Reference Distribution
Description
Uses root-finding to select a value of cens_adjust for
simulate_from_dgm such that a chosen censoring summary
statistic in the simulated data matches the corresponding statistic from
the DGM reference data (dgm$df_super).
Usage
calibrate_cens_adjust(
dgm,
target = c("rate", "km_median"),
n = 1000,
rand_ratio = 1,
analysis_time = 48,
max_entry = 24,
seed = 42,
interval = c(-3, 3),
tol = 1e-04,
tol_rel = 2.5,
n_eval = 2000,
verbose = TRUE,
...
)
Arguments
dgm |
An |
target |
Character. Calibration target: |
n |
Integer. Sample size passed to |
rand_ratio |
Numeric. Randomisation ratio passed to
|
analysis_time |
Numeric. Calendar analysis time passed to
|
max_entry |
Numeric. Maximum staggered entry time passed to
|
seed |
Integer. Base random seed. Each evaluation of the objective
function uses this seed for reproducibility. Default |
interval |
Numeric vector of length 2. Search interval for
|
tol |
Numeric. Root-finding tolerance. Default |
tol_rel |
Numeric. Maximum tolerated relative error, in percent; the
call stops if the achieved metric is not within |
n_eval |
Integer. Sample size used inside the objective function
during root-finding. Smaller values are faster but noisier; increase
for precision. Default |
verbose |
Logical. Print search progress and final result.
Default |
... |
Additional arguments passed to |
Details
Two calibration targets are supported:
"rate"Overall censoring rate (proportion censored). Finds
cens_adjustsuch thatmean(event_sim == 0)in simulated data equalsmean(event == 0)indgm$df_super."km_median"KM-based median censoring time, estimated by reversing the event indicator so censored observations become the "event" of interest. Finds
cens_adjustsuch that the simulated KM median matches the reference KM median.
How the objective function works
At each candidate cens_adjust value, the objective function:
Calls
simulate_from_dgm()withn = n_evaland the candidatecens_adjust.Calls
check_censoring_dgm()withverbose = FALSEto extract the target metric.Returns
sim_metric - ref_metric.
uniroot finds the zero crossing, i.e. the cens_adjust at
which simulated and reference metrics are equal.
Monotonicity
The objective is monotone in cens_adjust for both targets:
Larger
cens_adjust→ longer censoring times → lower censoring rate and higher KM median.Smaller
cens_adjust→ shorter censoring times → higher censoring rate and lower KM median.
If uniroot fails (the target lies outside the search interval),
the boundary values are printed and a wider interval should be
tried.
Stochastic noise
Because the objective function involves simulation, there is Monte Carlo
noise. Setting a fixed seed and a sufficiently large n_eval
(>= 2000) reduces noise enough for reliable root-finding. The
tol argument controls the root-finding tolerance on the
cens_adjust scale (not the metric scale).
Value
A named list with elements:
cens_adjustCalibrated
cens_adjustvalue.targetCalibration target used.
ref_valueReference metric value from
dgm$df_super.sim_valueAchieved metric value in simulated data at the calibrated
cens_adjust.residualAbsolute difference between
sim_valueandref_value.iterationsNumber of
unirootiterations.diagnosticOutput of
check_censoring_dgmat the calibrated value (invisibly).
See Also
simulate_from_dgm, check_censoring_dgm,
generate_aft_dgm_flex
Examples
## Not run:
library(survival)
# Build DGM on months scale
gbsg$time_months <- gbsg$rfstime / 30.4375
dgm <- generate_aft_dgm_flex(
data = gbsg,
continuous_vars = c("age", "size", "nodes", "pgr", "er"),
factor_vars = c("meno", "grade"),
outcome_var = "time_months",
event_var = "status",
treatment_var = "hormon",
subgroup_vars = c("er", "meno"),
subgroup_cuts = list(er = 20, meno = 0)
)
# Calibrate so simulated censoring rate matches reference
cal_rate <- calibrate_cens_adjust(
dgm = dgm,
target = "rate",
n = 1000,
analysis_time = 84,
max_entry = 24
)
cat("Calibrated cens_adjust (rate):", cal_rate$cens_adjust, "\n")
# Calibrate to KM median censoring time instead
cal_km <- calibrate_cens_adjust(
dgm = dgm,
target = "km_median",
n = 1000,
analysis_time = 84,
max_entry = 24
)
cat("Calibrated cens_adjust (km_median):", cal_km$cens_adjust, "\n")
# Use calibrated value in simulation
sim <- simulate_from_dgm(
dgm = dgm,
n = 1000,
analysis_time = 84,
max_entry = 24,
cens_adjust = cal_rate$cens_adjust,
seed = 123
)
mean(sim$event_sim) # event rate
mean(sim$event_sim == 0) # censoring rate — should match ref
## End(Not run)
Calibrate GLM Interaction for a Target Subgroup Effect Size
Description
Searches over a grid of interaction multipliers (k_inter) to find
the value that produces a target effect size in the subgroup Q.
Returns a modified "glm_dgm" object with the calibrated
interaction and updated true effects.
Usage
calibrate_glm_interaction(
data,
factor_vars,
continuous_vars = NULL,
outcome_var,
treatment_var,
target_effect,
outcome_type = c("binary", "continuous", "count"),
effect_measure = NULL,
offset_var = NULL,
subgroup_vars = NULL,
subgroup_cuts = NULL,
k_treat = 1,
adverse_outcome = FALSE,
k_inter_range = c(-10, 10),
grid_step = 0.05,
tol_rel = 2.5,
n_super = 5000L,
seed = 8316951L,
verbose = FALSE
)
Arguments
data |
The source data frame (same as passed to
|
factor_vars |
Character vector of factor variable names. |
continuous_vars |
Character vector of continuous prognostic variable
names. Passed through to |
outcome_var |
Character string naming the outcome variable. |
treatment_var |
Character string naming the treatment variable. |
target_effect |
Numeric. Target effect size in Q on the scale
determined by |
outcome_type |
Character. One of |
effect_measure |
Character. Effect measure. Default |
offset_var |
Character or |
subgroup_vars |
Character vector of subgroup-defining variables. |
subgroup_cuts |
Named list of cutpoint specifications. |
k_treat |
Numeric. Treatment effect scaling factor passed to
|
adverse_outcome |
Logical. When |
k_inter_range |
Numeric vector of length 2. Search range for
|
grid_step |
Numeric. Grid resolution. Default |
tol_rel |
Numeric. Maximum tolerated relative error, in percent; the
call stops (fails loudly) if the achieved OR is not within |
n_super |
Integer. Super-population size. Default |
seed |
Integer. Random seed. Default |
verbose |
Logical. Print calibration progress. Default |
Details
The function constructs a grid of k_inter values, calls
generate_glm_dgm for each, and selects the value
whose subgroup effect is closest to target_effect.
For binary outcomes with effect_measure = "OR", the target
is on the switched-treatment OR scale (OR > 1 = treatment increases
the outcome).
Value
An object of class "glm_dgm" with the calibrated
k_inter and updated hazard_ratios.
See Also
generate_glm_dgm,
simulate_from_glm_dgm
Examples
## Not run:
dgm <- calibrate_glm_interaction(
data = actg_df,
factor_vars = paste0("z", 1:12),
outcome_var = "y_binary",
treatment_var = "treat",
target_effect = 2.0,
outcome_type = "binary",
subgroup_vars = c("z1", "z2"),
subgroup_cuts = list(z1 = 1L, z2 = 1L),
verbose = TRUE
)
print(dgm)
## End(Not run)
Calibrate k_inter for Target Subgroup Hazard Ratio
Description
Finds the interaction effect multiplier (k_inter) that achieves a target hazard ratio in the harm subgroup.
Usage
calibrate_k_inter(
target_hr_harm,
model = "alt",
k_treat = 1,
cens_type = "weibull",
k_inter_range = c(-100, 100),
tol = 1e-06,
tol_rel = 2.5,
use_ahr = FALSE,
verbose = FALSE,
...
)
Arguments
target_hr_harm |
Numeric. Target hazard ratio for the harm subgroup |
model |
Character. Model type ("alt" only). Default: "alt" |
k_treat |
Numeric. Treatment effect multiplier. Default: 1 |
cens_type |
Character. Censoring type. Default: "weibull" |
k_inter_range |
Numeric vector of length 2. Search range for k_inter. Default: c(-100, 100) |
tol |
Numeric. Tolerance for root finding. Default: 1e-6 |
tol_rel |
Numeric. Maximum tolerated relative error, in percent; the
call stops if the achieved HR is not within |
use_ahr |
Logical. If TRUE, calibrate to AHR instead of Cox-based HR. Default: FALSE |
verbose |
Logical. Print diagnostic information. Default: FALSE |
... |
Additional arguments passed to |
Details
This function uses uniroot to find the k_inter value such that
the empirical HR (or AHR) in the harm subgroup equals target_hr_harm.
Value
Numeric value of k_inter that achieves the target HR
Examples
## Not run:
# Find k_inter for HR = 1.5 in harm subgroup
k <- calibrate_k_inter(target_hr_harm = 1.5, verbose = TRUE)
# Verify
dgm <- create_gbsg_dgm(model = "alt", k_inter = k, verbose = TRUE)
print(dgm$hr_H_true) # Should be close to 1.5
# Calibrate to AHR instead
k_ahr <- calibrate_k_inter(target_hr_harm = 1.5, use_ahr = TRUE, verbose = TRUE)
dgm_ahr <- create_gbsg_dgm(model = "alt", k_inter = k_ahr, verbose = TRUE)
print(dgm_ahr$AHR_H_true) # Should be close to 1.5
## End(Not run)
Calibrate k_treat for Target Overall Hazard Ratio
Description
Finds the treatment effect multiplier (k_treat) that achieves a
target overall hazard ratio in a DGM built by
generate_aft_dgm_flex. The root is found by uniroot
over k_treat with every other argument to
generate_aft_dgm_flex() held fixed.
Usage
calibrate_k_treat(
target_hr_overall,
base_args,
k_treat_range = c(-5, 5),
tol = 1e-06,
tol_rel = 2.5,
use_ahr = FALSE,
verbose = FALSE
)
Arguments
target_hr_overall |
Numeric. Target overall hazard ratio (must be positive). |
base_args |
Named list of arguments to
|
k_treat_range |
Numeric vector of length 2. Search range for
|
tol |
Numeric. Tolerance for root finding. Default |
tol_rel |
Numeric. Maximum tolerated relative error, in percent; the
call stops if the achieved HR is not within |
use_ahr |
Logical. If |
verbose |
Logical. Print diagnostic information. Default
|
Details
This is the natural sibling of calibrate_k_inter (which
targets the harm-subgroup HR under model = "alt") and
calibrate_cens_adjust (which targets the censoring rate).
It is typically used with model = "null" to fix the overall HR
for a uniform-effect DGM, but also works with model = "alt"
when the marginal HR is the calibration target rather than the
subgroup-specific HR.
Reproducibility requires the uniroot objective to be deterministic
in k_treat. This is achieved by including an integer
seed field in base_args so that every call to
generate_aft_dgm_flex() during the search uses the same
random-number stream. If base_args$seed is NULL the
objective becomes stochastic and uniroot may not converge; the
function emits a warning in this case.
Value
Numeric scalar. The calibrated k_treat value. Stops with an
error if the target cannot be reached – the search interval fails to
bracket it even after automatic extension, or the achieved HR lies outside
tol_rel percent of the target.
See Also
calibrate_k_inter,
calibrate_cens_adjust,
generate_aft_dgm_flex
Examples
## Not run:
library(survival)
data(gbsg)
gbsg$time_months <- gbsg$rfstime / 30.4375
base_args <- list(
data = gbsg,
continuous_vars = c("age", "size", "nodes", "pgr", "er"),
factor_vars = c("meno", "grade"),
outcome_var = "time_months",
event_var = "status",
treatment_var = "hormon",
model = "null",
n_super = 5000,
seed = 99,
verbose = FALSE
)
# Calibrate k_treat so that the overall HR is exactly 0.70
k <- calibrate_k_treat(target_hr_overall = 0.70,
base_args = base_args,
verbose = TRUE)
# Verify
dgm <- do.call(generate_aft_dgm_flex,
c(base_args, list(k_treat = k)))
print(dgm$hazard_ratios$overall) # ~ 0.70
# Calibrate to AHR instead
k_ahr <- calibrate_k_treat(target_hr_overall = 0.70,
base_args = base_args,
use_ahr = TRUE)
dgm_ahr <- do.call(generate_aft_dgm_flex,
c(base_args, list(k_treat = k_ahr)))
print(dgm_ahr$hazard_ratios$AHR) # ~ 0.70
## End(Not run)
Diagnose Censoring Consistency Between DGM Source Data and Simulated Data
Description
Compares the censoring distribution observed in the data used to build the
DGM against the censoring generated by simulate_from_dgm.
Reports censoring rates, time quantiles, KM-based median censoring times,
and flags substantial discrepancies.
Usage
check_censoring_dgm(
sim_data,
dgm,
treat_var = "treat_sim",
rate_tol = 0.1,
median_tol = 0.25,
verbose = TRUE
)
Arguments
sim_data |
A |
dgm |
An |
treat_var |
Character. Name of the treatment column in
|
rate_tol |
Numeric. Absolute tolerance (proportion scale) for
flagging a censoring-rate discrepancy. Default |
median_tol |
Numeric. Relative tolerance for flagging a KM median
censoring-time discrepancy. Default |
verbose |
Logical. If |
Details
The reference censoring distribution is derived from dgm$df_super,
sampled with replacement from the data passed to
generate_aft_dgm_flex(). Columns y (observed time) and
event (event indicator) in df_super reflect the original
observed censoring process on the DGM time scale.
The KM median censoring time is estimated by reversing the event indicator
(1 - event), treating events as censored and censored observations
as the event of interest. This gives a non-parametric estimate of the
censoring time distribution unconfounded by event occurrence.
Common causes of discrepancy: (1) time-scale mismatch (DGM built on days,
analysis_time in months); check exp(dgm$model_params$mu)
against your analysis_time. (2) Large cens_adjust shifting
censoring substantially from the fitted model. (3) Short
analysis_time or time_eos making administrative censoring
dominate the censoring process.
Value
Invisibly returns a named list. Elements are: rates (data
frame of censoring rates overall and by arm); quantiles (data
frame of censoring-time quantiles among censored subjects);
km_medians (data frame of KM-based median censoring times); and
flags (character vector of triggered warnings, empty if none).
See Also
simulate_from_dgm, generate_aft_dgm_flex
Examples
## Not run:
dgm <- setup_gbsg_dgm(model = "null", verbose = FALSE)
sim_data <- simulate_from_dgm(dgm, n = 200)
check_censoring_dgm(sim_data, dgm = dgm)
## End(Not run)
Confidence Interval for Estimate
Description
Calculates confidence interval for an estimate, optionally on log(HR) scale.
Usage
ci_est(x, sd, alpha = 0.025, scale = "hr", est.loghr = TRUE)
Arguments
x |
Numeric estimate. |
sd |
Numeric standard deviation. |
alpha |
Numeric significance level (default: 0.025). |
scale |
Character. "hr" or "1/hr". |
est.loghr |
Logical. Is estimate on log(HR) scale? |
Value
List with length, lower, upper, sd, and estimate.
Format a numeric threshold for a cut expression
Description
Renders a (already-rounded) numeric threshold as a clean string with no trailing zeros or floating-point noise, for use in "var <= value" cut expressions. Integer-valued thresholds print without a decimal point.
Usage
collapse_cuts_fmt(v, digits = 0L)
Arguments
v |
Numeric scalar. |
digits |
Integer. Decimal places used during rounding. |
Value
Character scalar.
Round half away from zero
Description
Base R round() uses round-half-to-even ("banker's rounding"), so
round(0.5) == 0 and round(74.5) == 74. Cut coarsening rounds
representative thresholds half up (away from zero) so that, e.g., a cluster
centroid of 74.5 maps to 75, matching the intuitive "nearest integer" rule.
Usage
collapse_cuts_round(x, digits = 0L)
Arguments
x |
Numeric vector. |
digits |
Integer. Number of decimal places. |
Value
Numeric vector rounded half away from zero.
Collapse near-redundant continuous candidate cuts
Description
Continuous covariates can generate many candidate threshold cuts that are practically redundant – e.g. "age <= 35.7" and "age <= 35", or "wtkg <= 75.2" and "wtkg <= 74.5" – because the candidate pool unions quantile cuts, GRF / DINA splits, and calibrated-DGM cuts at different precisions. This helper merges such near-duplicates to a single representative threshold, leaving categorical and indicator cuts untouched.
Usage
collapse_redundant_cuts(
cuts,
df,
confounders.name = NULL,
c_band = 1,
safety_tol = 0.05,
digits = 0L,
cont.cutoff = 4,
details = FALSE,
protect = character(0)
)
Arguments
cuts |
Character vector of candidate cut expressions, e.g.
|
df |
Data frame supplying the covariate columns, used for sd(x), the sample size n, and the membership safety check. |
confounders.name |
Character vector of confounder names (currently informational; variable identity is parsed from each cut expression). |
c_band |
Numeric >= 0. Band multiplier; |
safety_tol |
Numeric > 0. Membership safety tolerance: a fraction of n when < 1, an absolute subject count when >= 1. Default 0.05. |
digits |
Integer >= 0. Decimal places for the representative threshold. Default 0 (nearest integer). Raise for variables measured on a sub-unit scale. |
cont.cutoff |
Integer. A variable with at least this many unique values
is treated as continuous (see |
details |
Logical. If TRUE, print a per-cluster collapse report. |
protect |
Character vector of cut expressions to exempt from
coarsening (matched verbatim against |
Details
For each (variable, operator) group of continuous-variable cuts, distinct
thresholds are single-linkage clustered using a per-variable band
band = c_band * sd(x) / sqrt(n) (the standard error of the variable:
thresholds finer than ~1 SE are not statistically resolvable). Each cluster
collapses to one representative, the cluster mean rounded half-up to
digits places (see collapse_cuts_round()). Rounding alone can
also merge singletons that round to the same value (e.g. 75.2 and 74.5 both
round to 75 at digits = 0); de-duplication then drops the duplicate.
Safety check: a cluster is collapsed only if replacing each member threshold
by the representative changes subgroup membership for no more than
safety_tol subjects (a fraction of n when safety_tol < 1, an
absolute count when safety_tol >= 1). Clusters that would move more
than that are kept unchanged, so coarsening never silently redefines a
candidate subgroup by more than the tolerance.
Cut expressions that are categorical, indicator-valued (fewer than
cont.cutoff unique values), bare variable names, or equality tests
(==) pass through unchanged. Cuts listed in protect are also
passed through verbatim: they are model-identified screening candidates
(e.g. GRF/DINA cuts) rather than points on a default continuous grid, so
coarsening them would defeat the purpose of the screening engine.
Value
Character vector of de-duplicated cut expressions (length <= input).
Collect Results from foreach with Error Handling
Description
Separates successful results from errors in the output of a
foreach loop run with
.errorhandling = "pass". Successful results are combined via
rbindlist (with an rbind fallback
if data.table is not available) and failed tasks are counted,
reported, and optionally escalated to a hard error.
Usage
collect_results(
raw_list,
label = "",
stop_on_all_fail = TRUE,
max_error_messages = 3L,
keep_diagnostics = FALSE
)
Arguments
raw_list |
List returned by |
label |
Character. Optional label included in the warning
message to identify which simulation phase produced the errors
(e.g. |
stop_on_all_fail |
Logical. If |
max_error_messages |
Integer. Maximum number of distinct
error messages to print in the diagnostic output. Default
|
keep_diagnostics |
Logical. If |
Details
This replaces the fragile pattern of calling nrow() or
rbind() directly on a raw foreach result list, which
fails silently when error objects are mixed in with data-frame rows.
Design rationale for the reporting strategy:
A
warningis emitted so that calling code (andtryCatch/withCallingHandlersblocks) can capture the failure condition programmatically.The error summary is also printed with
cattostderr, becausewarningandmessageoutput is sometimes buffered or suppressed inside future/foreach parallel workers or whenoptions(warn = -1)is set by user code upstream. Thecat()output is the visible-no-matter-what fallback.When
stop_on_all_fail = TRUE(the default), a completely empty result list is escalated to an error rather than silently returning an empty data frame. This avoids downstream code proceeding with zero rows and producing misleading summaries. Calibration code that may legitimately tolerate some failures (e.g. Monte Carlo FPR estimation) can setstop_on_all_fail = FALSE.
The returned object is always a data frame – never NULL –
because attributes (n_failed, n_total) attached to
NULL are silently discarded by R, whereas attributes on an
empty data frame are preserved and inspectable by the caller.
Value
A data frame of row-bound successful results. Always a
data frame; never NULL – an empty data frame is
returned when every task failed and stop_on_all_fail = FALSE.
Two attributes are always attached:
n_failedInteger. Number of tasks that returned errors.
n_totalInteger. Total number of tasks (successes + failures).
When keep_diagnostics = TRUE, a third attribute
"diagnostics" may be attached – a named list of
per-replicate diagnostic objects, keyed by "sim<ID>"
when the source element carries an "sim_id" attribute
(run_simulation_analysis sets this automatically).
Preserving per-replicate diagnostics
data.table::rbindlist() drops per-element attributes during
row-binding. When the caller has populated each replicate with
heavy diagnostic objects (e.g. via
run_simulation_analysis with a non-empty
keep argument), set keep_diagnostics = TRUE to
retain them. The diagnostics attribute on each input element
must carry attr(..., "sim_id") for stable keying; that is
set automatically by run_simulation_analysis().
See Also
reset_workers for resetting the
future plan and releasing worker memory between
simulation phases.
Examples
library(foreach)
library(doFuture)
future::plan("sequential")
raw <- foreach(i = 1:5, .errorhandling = "pass") %dofuture% {
if (i == 3L) stop("simulated failure")
data.frame(task = i, value = rnorm(1))
}
res <- collect_results(raw, label = "demo")
# Warning: demo: 1 of 5 tasks failed.
nrow(res)
# [1] 4
attr(res, "n_failed")
# [1] 1
Compare Detection Curves Across Sample Sizes
Description
Generates and compares detection probability curves for multiple subgroup sample sizes.
Usage
compare_detection_curves(
n_sg_values,
prop_cens = 0.3,
hr_threshold = 1.25,
hr_consistency = 1,
theta_range = c(0.5, 3),
n_points = 40L,
verbose = TRUE
)
Arguments
n_sg_values |
Integer vector. Subgroup sample sizes to compare. |
prop_cens |
Numeric. Proportion censored. Default: 0.3 |
hr_threshold |
Numeric. HR threshold. Default: 1.25 |
hr_consistency |
Numeric. HR consistency threshold. Default: 1.0 |
theta_range |
Numeric vector of length 2. Range of HR values. Default: c(0.5, 3.0) |
n_points |
Integer. Number of points per curve. Default: 40 |
verbose |
Logical. Print progress. Default: TRUE |
Value
A data.frame with all curves combined, including n_sg as a factor.
Examples
## Not run:
comparison <- compare_detection_curves(
n_sg_values = c(40, 60, 80, 100),
prop_cens = 0.2
)
# Plot with ggplot2
library(ggplot2)
ggplot(comparison, aes(x = theta, y = probability, color = factor(n_sg))) +
geom_line(linewidth = 1) +
labs(x = "True HR", y = "P(Detect)", color = "n_sg") +
theme_minimal()
## End(Not run)
Compare Subgroup Membership Across Two Analyses
Description
Cross-tabulates subgroup assignments from two analyses and produces
a formatted gt summary table with concordance statistics.
Accepts forestsearch result objects, GRF result lists, or raw
indicator vectors. Non-detection (no subgroup found) is handled
automatically: all subjects are classified to the complement.
Usage
compare_fs_subgroups(
fs1,
fs2,
label1 = "Analysis 1",
label2 = "Analysis 2",
id.name = NULL,
sg0_label = NULL,
sg1_label = NULL,
subgroup_notation = c("harm", "benefit"),
title = "Subgroup Membership Comparison",
subtitle = NULL,
font_size = 12
)
Arguments
fs1 |
Either a |
fs2 |
Same types as |
label1 |
Character. Display label for the first analysis. |
label2 |
Character. Display label for the second analysis. |
id.name |
Character. Subject ID column for merge alignment.
Only used when both inputs are result objects.
If |
sg0_label |
Character or |
sg1_label |
Character or |
subgroup_notation |
Character. |
title |
Character. Table title. |
subtitle |
Character or |
font_size |
Numeric. Font size for gt table. |
Value
A list with:
- table
A
gtcross-tabulation table.- crosstab
The raw cross-tabulation matrix.
- concordance
Agreement, kappa, counts.
- membership
Per-subject assignments from both analyses.
Examples
## Not run:
# Harm search: Cox vs Poisson (both forestsearch objects)
result <- compare_fs_subgroups(fs_cox, fs_pois,
label1 = "Cox PH", label2 = "Poisson",
subgroup_notation = "harm")
# Benefit search: FS vs GRF (handles non-detection automatically)
result <- compare_fs_subgroups(fs_result, grf_result,
label1 = "ForestSearch", label2 = "GRF",
subgroup_notation = "benefit")
# Raw vectors
result <- compare_fs_subgroups(
c(rep(0L, 30), rep(1L, 70)),
c(rep(0L, 25), rep(1L, 75)),
label1 = "Method A", label2 = "Method B")
## End(Not run)
Compare Multiple Survival Regression Models
Description
Performs comprehensive comparison of multiple survreg models including convergence checking, information criteria comparison, and model selection.
Usage
compare_multiple_survreg(
...,
model_names = NULL,
verbose = TRUE,
criteria = c("AIC", "BIC")
)
Arguments
... |
survreg model objects to compare |
model_names |
Optional character vector of model names |
verbose |
Logical, whether to print detailed output (default: TRUE) |
criteria |
Character vector of criteria to use ("AIC", "BIC", or both) |
Value
A list of class "multi_survreg_comparison" containing:
- models
Named list of input models
- convergence
Convergence status for each model
- comparison
Model comparison statistics
- rankings
Model rankings by different criteria
- best_model
Name of the best model
- recommendation
Text recommendation
Examples
## Not run:
fit1 <- survreg(Surv(time, status) ~ x, dist = "weibull")
fit2 <- survreg(Surv(time, status) ~ x, dist = "lognormal")
comparison <- compare_multiple_survreg(fit1, fit2)
## End(Not run)
Compare forestsearch Runs Across Selection-Rule Combinations
Description
Runs forestsearch once per combination of
sg_focus[i] + selection_rule[i] (tuple semantics; the
two vectors are paired element-by-element and must have the same
length). All other arguments are held fixed across runs and passed
through via ....
Usage
compare_selection_rules(
df.analysis,
sg_focus,
selection_rule,
compute_cis = TRUE,
n_splits = 1000L,
ci_seed = 1L,
plot_xlim = NULL,
show_band = TRUE,
combo_labels = NULL,
verbose = TRUE,
...
)
Arguments
df.analysis |
Analysis-ready data frame; passed to
|
sg_focus |
Character vector of |
selection_rule |
Character vector of |
compute_cis |
Logical. If |
n_splits |
Integer. Passed to
|
ci_seed |
Integer. Passed to
|
plot_xlim |
Numeric vector of length 2 (or |
show_band |
Logical. Whether to draw the effect-band
shading on each plot. Default |
combo_labels |
Character vector of length |
verbose |
Logical. Print one line per combination during
execution. Default |
... |
Additional arguments passed through to every
|
Details
For each combination the wrapper:
Captures the in-flight console output from forestsearch (including the pre-consistency Candidate Evaluation Preview and post-consistency Candidate Evaluation Summary when
show_candidate_summary = TRUEis passed).Optionally computes the frontier CIs via
compute_frontier_cis.Builds the Pareto-frontier plot via
plot_pareto_frontier.
The captured stdout is returned per combination so the caller can
cat() it later for side-by-side inspection in a Quarto chunk.
Plot objects are returned individually and (when
patchwork is installed) composed into a single side-by-side
figure. The full fs object for each combination is returned
so downstream diagnostics (pareto_frontier_table,
explain_pareto_selection) can be run without re-fitting.
Value
A list with class "forestsearch_comparison" and
components:
combosdata.framewith columnssg_focus,selection_rule,label.fsNamed list of
forestsearchresult objects, one per combination. NULL entries mark fits that errored.ci_tabNamed list of frontier-CI tables (or all NULL if
compute_cis = FALSE).plotsNamed list of
ggplotobjects, one per combination.plot_gridA single patchwork object placing the plots side by side, or
NULLif patchwork isn't installed.plot_combinedA single
ggplotcomposing all configurations onto one Pareto plot, with each winner labeledS1: <combo_label>,S2: <combo_label>, etc. Returned only when ALL successful configurations share an identical passing set (seeplot_pareto_combinedfor the equality criterion);NULLotherwise. Seeplot_combined_subsetsbelow for the per-subset fallback when some – but not all – configurations match.plot_combined_subsetsNamed list of
ggplotobjects, one per equivalence class of configurations sharing an identical passing set. Only groups of size\ge 2are included (a singleton group has no peer to combine with). Names are derived from the sharedsg_focuswhen constant within a group (e.g.,"effMaxSG"); from the concatenated combo labels otherwise. When all valid combos share one passing set this list contains exactly one element, equal toplot_combined. The list is empty when no two combos match.combined_skip_reasonCharacter scalar or
NULL. Whenplot_combinedisNULLbecause the equality precondition failed, this captures the specific reason – size mismatch, definition-set mismatch, or value drift on a named column – as a single string.NULLwhen the combined plot was built or when fewer than two fits succeeded.consoleNamed list of character vectors – the captured stdout from each forestsearch run (full output).
diagnosticsNamed list of per-combo slices of the captured output: each element is a list with character-scalar fields
preview(theCANDIDATE EVALUATION PREVIEWblock),summary(theCANDIDATE EVALUATION SUMMARYblock), andfull(the entire captured stdout, same content asconsole[[i]]collapsed to a single string).previewandsummaryareNA_character_when the corresponding banner is absent (e.g.\show_candidate_summarywasFALSE, or the run errored before reaching it). This slice is what a comparison document should display as the primary diagnostic;full(orconsole[[i]]) remains available for an expandable / on-demand view of the complete run output.errorsNamed list of error messages (or NULL) for combos that failed.
Tuple semantics
With sg_focus = c("hrMaxSG", "hrMaxSG") and
selection_rule = c("pareto", "both"), two runs are
performed: one with (hrMaxSG, pareto) and one with
(hrMaxSG, both). The length-mismatch case raises an error.
Console capture
All stdout from each forestsearch() call is captured (not
just the diagnostic tables) via capture.output(). The
PREVIEW / SUMMARY blocks are clearly delineated by banner separators
within the captured text. Use cat(out$console[[i]]) to
replay the output in a Quarto chunk; results = "asis" works
but is not required since the captured text already includes
newlines.
See Also
forestsearch,
plot_pareto_frontier,
plot_pareto_combined,
pareto_frontier_table,
explain_pareto_selection,
compute_frontier_cis.
Examples
## Not run:
out <- compare_selection_rules(
df.analysis = actg_df,
sg_focus = c("hrMaxSG", "hrMaxSG"),
selection_rule = c("pareto", "both"),
# All other forestsearch args:
confounders.name = confounders.name,
outcome.name = adverse_outcome,
treat.name = treat.name,
id.name = id.name,
outcome_type = "continuous",
effect_measure = "MD",
adverse_outcome = TRUE,
hr.threshold = 10,
hr.consistency = 5,
pconsistency.threshold = 0.90,
show_candidate_summary = TRUE,
details = TRUE
)
# Inspect captured console output for combo 1
cat(out$console[[1]])
# Side-by-side plot
print(out$plot_grid)
# Use fs object for downstream diagnostics
pareto_frontier_table(out$fs[[1]], ci_table = out$ci_tab[[1]])
explain_pareto_selection(out$fs[[2]], ci_table = out$ci_tab[[2]])
## End(Not run)
Compare two extreme-subgroups simulation studies
Description
Builds the per-subgroup comparison frame between two
run_subgroup_sims() results (or RDS payloads loaded from the
vignettes – any list carrying sim_hrs / sim_ubs / sim_ns
matrices with identical column names). For each input the per-design
statistics replicate the vignettes' Section 6.6.1 definitions
exactly: tail probabilities are percentages denominated over
converged fits (na.rm = TRUE), medians are unconditional, N is
the across-trial mean subgroup size, and structurally empty
subgroups (all-NA columns) are mapped from NaN to NA.
Usage
compare_subgroup_sims(
x,
y,
expect_designs = NULL,
suffixes = c("_r", "_f"),
est_thresholds = NULL,
ub_thresholds = NULL
)
Arguments
x, y |
The two studies to compare, in display order (the memo passes random-X then fixed-X). |
expect_designs |
Optional length-2 character: required |
suffixes |
Length-2 character appended to each statistic's
column for |
est_thresholds |
Length-2 numeric |
ub_thresholds |
Length-2 numeric for the |
Details
Column names are structural and retained across outcome types
(hr05, hr1, ub2, ub3, mHR, mUB), exactly as sim_hrs
serves generic duty on GLM results: for subgroup_glm() fits they
hold estimate-scale statistics at the resolved thresholds. Threshold
resolution: explicit arguments win; otherwise the inputs' effect
metadata supplies them; otherwise the HR legacy c(0.5, 1.0) /
c(2, 3). An NA threshold yields an all-NA column. The two
inputs must carry compatible metadata (both none, or the same
effect measure) – comparing an MD study against an HR study is a
hard error, raised before the panel-alignment guard so the
categorical incompatibility is reported first.
Two validations guard the row-wise alignment (formerly the memo's
Guard 1 and Guard 2): the subgroup panels must be identical, and,
when expect_designs is supplied, each input's design label must
match it exactly.
The medians here use median(), matching the memo; they are equal
to summary.subgroup_sims()'s type-7 quantile(..., 0.50) values
up to floating-point formulation ((a + b) / 2 versus
a + 0.5 * (b - a)), which is why the memo's numbers and the
vignettes' tables have always agreed to the printed digit.
Value
A plain data.frame (one row per subgroup, matching the
memo's static-mode frame type) with columns subgroup, then for
each of N, ub2, ub3, mUB, hr05, hr1, mHR the x and
y columns interleaved. The inputs' n_sims and design values
are attached as attributes "n_sims" and "designs"; when the
inputs carry effect metadata it is attached as attribute
"effect" (informational; attributes are dropped by most
transformations).
See Also
run_subgroup_sims(), summary.subgroup_sims(),
subgroup_glm()
Examples
## Not run:
pl_r <- readRDS("results/extreme_sims_resample_10000_payload.rds")
pl_f <- readRDS("results/extreme_sims_fixed_10000_payload.rds")
cmp <- compare_subgroup_sims(pl_r, pl_f,
expect_designs = c("resample", "fixed"))
## End(Not run)
Compute AHR from loghr_po
Description
Computes Average Hazard Ratio from individual log hazard ratios.
Usage
compute_ahr(df, subset_indicator = NULL)
Arguments
df |
Data frame with loghr_po column |
subset_indicator |
Optional logical/integer vector for subsetting |
Value
Numeric AHR value
Compute Marginal Causal Effect from GLM Potential Outcomes
Description
Computes the marginal (average) treatment effect from individual-level
potential outcomes stored in GLM simulation data.
This is the GLM analogue of compute_ahr for survival data.
Usage
compute_aor(df, subset_indicator = NULL, effect_measure = "OR")
Arguments
df |
Data frame with potential outcome columns ( |
subset_indicator |
Optional logical/integer vector for subsetting.
If provided, only rows where |
effect_measure |
Character. One of |
Details
For binary outcomes (p0, p1):
-
OR:odds(mean(p1)) / odds(mean(p0)) -
RD:mean(p1) - mean(p0) -
RR:mean(p1) / mean(p0)
For continuous or count outcomes (mu0, mu1):
-
MD:mean(mu1) - mean(mu0) -
IRR:mean(mu1) / mean(mu0) -
IRD:mean(mu1) - mean(mu0)
Value
Numeric effect value, or NA_real_ if required columns
are missing.
Compute CDE from theta_0 and theta_1
Description
Computes Controlled Direct Effect as the ratio of average hazard
contributions on the natural scale:
CDE(S) = mean(exp(theta_1[S])) / mean(exp(theta_0[S])).
Usage
compute_cde(df, subset_indicator = NULL)
Arguments
df |
Data frame with |
subset_indicator |
Optional logical/integer vector for subsetting.
If provided, only rows where |
Value
Numeric CDE value, or NA_real_ if columns are missing.
Compute CDE Analogue from GLM Potential Outcomes
Description
Computes the Controlled Direct Effect analogue for GLM outcomes.
For binary (logit link), this averages the individual-level odds
first, then takes the ratio — analogous to
mean(exp(theta_1)) / mean(exp(theta_0)) in survival:
Usage
compute_cde_glm(df, subset_indicator = NULL, effect_measure = "OR")
Arguments
df |
Data frame with potential outcome columns ( |
subset_indicator |
Optional logical/integer vector for subsetting.
If provided, only rows where |
effect_measure |
Character. One of |
Details
CDE(S) = mean(p1[S]/(1-p1[S])) / mean(p0[S]/(1-p0[S]))
This differs from the AOR (odds(mean(p1)) / odds(mean(p0)))
due to Jensen's inequality whenever individual probabilities are
heterogeneous. For identity-link (MD) and log-link (IRR) outcomes,
CDE equals AOR and this function returns NA_real_.
Value
Numeric CDE value (binary OR only), or NA_real_.
Compute Probability of Detecting True Subgroup
Description
Calculates the probability that a true subgroup with given hazard ratio will be detected using the ForestSearch consistency-based criteria.
Usage
compute_detection_probability(
theta,
n_sg,
prop_cens = 0.3,
hr_threshold = 1.25,
hr_consistency = 1,
method = c("cubature", "monte_carlo"),
n_mc = 100000L,
tol = 1e-04,
verbose = FALSE
)
Arguments
theta |
Numeric. True hazard ratio in the subgroup. Can be a vector for computing detection probability across multiple HR values. |
n_sg |
Integer. Subgroup sample size. |
prop_cens |
Numeric. Proportion censored (0-1). Default: 0.3 |
hr_threshold |
Numeric. HR threshold for detection (e.g., 1.25). This is the threshold that the average HR across splits must exceed. |
hr_consistency |
Numeric. HR consistency threshold (e.g., 1.0). This is the threshold each individual split must exceed. Default: 1.0 |
method |
Character. Integration method: "cubature" (recommended for accuracy) or "monte_carlo" (faster for exploration). Default: "cubature" |
n_mc |
Integer. Number of Monte Carlo samples if method = "monte_carlo". Default: 100000 |
tol |
Numeric. Relative tolerance for cubature integration. Default: 1e-4 |
verbose |
Logical. Print progress for vector inputs. Default: FALSE |
Details
This function computes P(detect | theta) using the asymptotic normal approximation for the log hazard ratio estimator. The detection criterion is based on ForestSearch's split-sample consistency evaluation:
The subgroup HR estimate must exceed hr_threshold on average
Each split-half must individually exceed hr_consistency
The approximation assumes:
Large sample sizes (CLT applies)
Var(log(HR)) ~ 4/d for the full subgroup (d = total subgroup events); each 50/50 split-half (d/2 events) has variance 8/d
Independence between split-halves (conditional on true effect)
Value
If theta is scalar, returns a single probability. If theta is a vector, returns a data.frame with columns: theta, probability.
Examples
## Not run:
# Single HR value
prob <- compute_detection_probability(
theta = 1.5,
n_sg = 60,
prop_cens = 0.2,
hr_threshold = 1.25
)
# Vector of HR values for power curve
hr_values <- seq(1.0, 2.5, by = 0.1)
results <- compute_detection_probability(
theta = hr_values,
n_sg = 60,
prop_cens = 0.2,
hr_threshold = 1.25,
verbose = TRUE
)
# Plot detection probability curve
plot(results$theta, results$probability, type = "l",
xlab = "True HR", ylab = "P(detect)")
## End(Not run)
Compute Detection Probability for GLM Outcomes
Description
Extends the Section 2.1 approximation (Leon et al., 2024, eq. 3) to
arbitrary GLM outcomes via the effective information
d_{\text{eff}}. This is the GLM generalization of
compute_detection_probability.
Usage
compute_detection_probability_glm(
theta,
d_eff,
c1 = 1.25,
c2 = 1,
effect_scale = c("ratio", "difference"),
method = c("cubature", "monte_carlo"),
n_mc = 500000L,
seed = 42L,
verbose = FALSE
)
Arguments
theta |
Numeric (scalar or vector). True treatment effect in the subgroup on the natural scale:
|
d_eff |
Numeric. Effective information in the subgroup.
Use the |
c1 |
Numeric. Screening threshold on the natural scale. For ratio measures: the ratio (e.g., 1.25). For difference measures: the difference (e.g., 0.2). |
c2 |
Numeric. Consistency threshold on the natural scale. Default: 1.0 for ratio, 0.0 for difference. |
effect_scale |
Character. |
method |
Character. |
n_mc |
Integer. Monte Carlo samples if method = "monte_carlo". |
seed |
Integer. RNG seed for Monte Carlo. |
verbose |
Logical. Progress messages for vector theta. |
Details
The detection criterion (Leon et al., 2024, Section 2.1):
Screening:
W_1 + W_2 \ge 2k_1Consistency:
\min(W_1, W_2) \ge k_2
with W_1, W_2 \sim N(\beta, 8/d_{\text{eff}}) independently.
For ratio-scale effects, \beta = \log(\theta),
k_j = \log(c_j).
For difference-scale effects, \beta = \theta, k_j = c_j.
Value
If theta is scalar, a single probability.
If vector, a data.frame with columns theta, probability.
See Also
compute_detection_probability for the
survival-specific version.
Examples
# Survival: reproduce Figure 2
d <- d_eff_survival(n_sg = 60, prop_cens = 0.45)
compute_detection_probability_glm(theta = 2.0, d_eff = d)
# Binary: OR = 2.0, event rate 30%
d <- d_eff_binary(n_sg = 100, p_event = 0.30)
compute_detection_probability_glm(theta = 2.0, d_eff = d)
# Count: IRR = 2.0, 80 total events
d <- d_eff_count(total_events = 80)
compute_detection_probability_glm(theta = 2.0, d_eff = d)
# Continuous: mean difference = 0.5, sigma = 1.5
d <- d_eff_continuous(n_sg = 100, sigma_y = 1.5)
compute_detection_probability_glm(
theta = 0.5, d_eff = d,
c1 = 0.2, c2 = 0.0,
effect_scale = "difference"
)
# Power curve
d <- d_eff_binary(100, 0.30)
result <- compute_detection_probability_glm(
theta = seq(0.5, 3.0, by = 0.1), d_eff = d
)
plot(result$theta, result$probability, type = "l")
Compute Detection Probability for Single Theta (Internal)
Description
Compute Detection Probability for Single Theta (Internal)
Usage
compute_detection_probability_single(
theta,
n_sg,
prop_cens,
k_avg,
k_ind,
method,
n_mc,
tol
)
Arguments
theta |
Numeric. True hazard ratio in the subgroup. Can be a vector for computing detection probability across multiple HR values. |
n_sg |
Integer. Subgroup sample size. |
prop_cens |
Numeric. Proportion censored (0-1). Default: 0.3 |
k_avg |
Log of hr_threshold |
k_ind |
Log of hr_consistency |
method |
Character. Integration method: "cubature" (recommended for accuracy) or "monte_carlo" (faster for exploration). Default: "cubature" |
n_mc |
Integer. Number of Monte Carlo samples if method = "monte_carlo". Default: 100000 |
tol |
Numeric. Relative tolerance for cubature integration. Default: 1e-4 |
Value
Numeric probability
Compute and Attach CDE Values to a DGM Object
Description
Calculates Controlled Direct Effect (CDE) hazard ratios from the
super-population potential outcomes (theta_0, theta_1)
and attaches them to the DGM's hazard_ratios list. This enables
automatic CDE detection by build_estimation_table.
Usage
compute_dgm_cde(dgm, harm_col = NULL)
Arguments
dgm |
A DGM object (e.g., from |
harm_col |
Character. Name of the subgroup indicator column in
|
Details
The CDE for subgroup S is defined as:
CDE(S) = mean(exp(theta_1[S])) / mean(exp(theta_0[S]))
which is the ratio of average hazard contributions on the natural scale.
This differs from the AHR (exp(mean(loghr_po))) due to Jensen's
inequality. In the notation of Leon et al. (2024), CDE corresponds to
theta-ddagger.
The function detects the subgroup indicator column automatically,
checking for flag.harm, flag_harm, and H in
the super-population data frame.
Value
The DGM object with CDE values added to
dgm$hazard_ratios (CDE, CDE_harm,
CDE_no_harm) and to top-level fields (dgm$CDE,
dgm$cde_H, dgm$cde_Hc).
See Also
build_estimation_table, get_dgm_hr
Examples
## Not run:
dgm <- create_gbsg_dgm(model = "alt", k_inter = 2.0)
dgm <- compute_dgm_cde(dgm)
dgm$hazard_ratios$CDE_harm # theta-ddagger(H)
dgm$hazard_ratios$CDE # theta-ddagger overall
## End(Not run)
Compute Naive and Split-Derived 95% CIs for Pareto Frontier Members
Description
For each candidate subgroup on fs$grp.consistency$out_sg$pareto_frontier,
computes three 95\
Usage
compute_frontier_cis(
fs,
n_splits = 1000L,
ci_level = 0.95,
seed = NULL,
verbose = FALSE
)
Arguments
fs |
A |
n_splits |
Integer. Number of 50/50 random splits per frontier
member for the Split CI. Default |
ci_level |
Numeric in |
seed |
Integer or |
verbose |
Logical. Default |
Details
-
Naive CI: full-sample Wald CI from a Cox or GLM model refit on the subgroup's data. This is the textbook CI that ignores the multiple-testing / subgroup-search context. Anti-conservative by construction.
-
Split CI: a 95\ estimate, with SE estimated by the empirical standard deviation of the
2Sindividual half-sample effect estimates produced byn_splitsrandom 50/50 splits. This is a half-jackknife SE (Shao1996). Naming: the\simin column labels marks this as a resampling-derived approximation, not a model-based CI. -
FSBC-mimic CI: a bias-corrected interval inspired by the bootstrap bias-correction algorithm of Leon2024fs (eq 7), with the selection-on-bootstrap-data term
\eta_b^*(\hat H_b^*)zeroed out (the selected subgroup is treated as fixed across half-jackknife replicates). The bias-corrected estimate is\hat\beta^{\mathrm{FSBC}} = 2\hat\beta - \overline{\hat\beta^{(h)}}where\overline{\hat\beta^{(h)}}is the mean of the2Shalf-sample estimates; the SE is the same half-jackknife SE as for the Split CI. See “Details — FSBC-mimic interpretation” below.
All three CIs are computed post-hoc on the returned forestsearch
object; this function does not modify any internal state.
Value
A data.table keyed by frontier row index (m) with
columns:
mCandidate id (matches
pareto_frontier$m).estimateRefit effect estimate on the natural scale.
naive_lcl,naive_uclNaive Wald CI bounds.
split_sdHalf-jackknife SE estimate: empirical SD of the
2Sper-half effect estimates on the SE scale (log for ratio measures, natural for differences).split_lcl,split_uclSplit-derived CI bounds.
fsbc_estimateBias-corrected effect estimate on the natural scale (
2\hat\beta - \overline{\hat\beta^{(h)}}on the SE scale, then back-transformed).fsbc_lcl,fsbc_uclFSBC-mimic CI bounds.
n_valid_splitsNumber of splits that produced valid per-half estimates.
Rows where the refit fails (e.g., zero events in one arm) have
NA in all CI columns.
Details — Half-jackknife SE
For each of n_splits random 50/50 splits we fit the effect model
separately on each half, obtaining 2S per-half estimates
\{\hat\beta^{(s,h)} : s=1,\ldots,S; h=1,2\}. The half-jackknife
SE estimator is
\widehat{\mathrm{SE}} = \mathrm{sd}\bigl\{\hat\beta^{(s,h)}\bigr\},
i.e., the plain empirical SD of the 2S estimates. No
\sqrt{2} correction is applied: simulation studies confirm that
this estimator's CI achieves nominal coverage in a no-selection
setting (see the package's pressure-test script). Earlier versions
of this function used the average (\hat\beta^{(s,1)} +
\hat\beta^{(s,2)})/2 per split; that approach is invalid because
the two halves are complementary and their average is nearly
constant.
Details — FSBC-mimic interpretation
The FSBC-mimic CI is a quick approximation to the bootstrap
bias-corrected estimator of Leon2024fs. Substituting the
2S per-half estimates for bootstrap replicates, and treating
the selected subgroup \hat H as fixed (i.e., omitting the
\eta_b^*(\hat H_b^*) = \hat\beta_b^*(\hat H_b^*) -
\hat\beta(\hat H_b^*) term that captures variability in FS selection
on the bootstrap data), the bias-corrected estimator from paper
eq (7) collapses to
\hat\beta^{\mathrm{FSBC}}(\hat H) = \hat\beta(\hat H) -
\frac{1}{2S}\sum_{s,h}\bigl[\hat\beta^{(s,h)} - \hat\beta(\hat H)\bigr]
= 2\hat\beta(\hat H) - \overline{\hat\beta^{(h)}}.
The SE is the same half-jackknife SE as for the Split CI.
The FSBC-mimic CI is not a full FSBC CI: it omits the
selection-on-bootstrap-data variance. Use this column as a quick
bias-corrected diagnostic, not for inference. For full FSBC CIs use
forestsearch_bootstrap_dofuture.
See Also
pareto_frontier_table,
frontier_member_flags, plot_pareto_frontier.
Compute node metrics for a policy tree
Description
Aggregates scores by leaf node and calculates treatment effect differences
Usage
compute_node_metrics(data, dr.scores, tree, X, n.min)
Arguments
data |
Data frame. Original data |
dr.scores |
Matrix. Doubly robust scores |
tree |
Policy tree object |
X |
Matrix. Covariate matrix |
n.min |
Integer. Minimum subgroup size |
Value
Data frame with node metrics
Examples
## Not run:
# compute_node_metrics() is called internally by grf.subg.harm.survival().
# See grf.subg.harm.survival() for the standard entry point.
## End(Not run)
Compute Pareto Frontier on (Effect, N)
Description
Returns the subset of candidate subgroups that are not dominated on
the two-objective space (effect size, sample size), where both
objectives are maximized. A candidate i is dominated iff
another candidate j has hr_j \ge hr_i and
N_j \ge N_i, with at least one inequality strict.
Usage
compute_pareto_frontier(result_dt, effect_log_scale = FALSE)
Arguments
result_dt |
A data.table with numeric columns |
effect_log_scale |
Logical. If |
Details
The frontier is intended as a post-hoc reporting
artifact, not a selection criterion. The selected subgroup
(under sg_focus) may or may not appear on the frontier –
in particular, "hrMinSG" may select an N-dominated point
by design (preferring small subgroups).
Value
A data.table of non-dominated rows, sorted by hr
descending. Returns an empty 0-row data.table if input is empty.
Compute Hazard Ratio for a Single Subgroup
Description
Internal helper function to compute HR and CI for a subgroup. Uses robust (sandwich) standard errors for consistency with cox_summary().
Usage
compute_sg_hr(
df,
sg_name,
outcome.name,
event.name,
treat.name,
E.name,
C.name,
z_alpha = qnorm(0.975),
conf.level = 0.95
)
Arguments
df |
Data frame for the subgroup. |
sg_name |
Character. Name of the subgroup. |
outcome.name |
Character. Name of survival time variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
E.name |
Character. Label for experimental arm. |
C.name |
Character. Label for control arm. |
z_alpha |
Numeric. Z-multiplier for CI (default: qnorm(0.975) for 95% CI). |
conf.level |
Numeric. Confidence level for intervals (default: 0.95). |
Value
Data frame with single row of HR estimates, or NULL if model fails.
Compute Hazard Ratio Estimates for Subgroups
Description
Internal function to compute Cox model hazard ratio estimates with confidence intervals for ITT, H, and Hc subgroups.
Usage
compute_sg_hr_estimates(
df,
df_H,
df_Hc,
outcome.name,
event.name,
treat.name,
conf.level = 0.95,
verbose = FALSE
)
Arguments
df |
Full analysis data frame |
df_H |
Data frame for H subgroup |
df_Hc |
Data frame for Hc subgroup |
outcome.name |
Character. Outcome variable name |
event.name |
Character. Event indicator name |
treat.name |
Character. Treatment variable name |
conf.level |
Numeric. Confidence level |
verbose |
Logical. Print messages |
Value
Data frame with HR estimates
Compute Summary Statistics for Subgroups
Description
Internal function to compute summary statistics for each subgroup.
Usage
compute_sg_summary(
df,
df_H,
df_Hc,
outcome.name,
event.name,
treat.name,
sg0_name,
sg1_name
)
Arguments
df |
Full analysis data frame |
df_H |
Data frame for H subgroup |
df_Hc |
Data frame for Hc subgroup |
outcome.name |
Character. Outcome variable name |
event.name |
Character. Event indicator name |
treat.name |
Character. Treatment variable name |
sg0_name |
Character. Label for H subgroup |
sg1_name |
Character. Label for Hc subgroup |
Value
Data frame with summary statistics
Wald confidence intervals for DINA coefficients.
Description
Computes normal-approximation confidence intervals
\hat\beta_j \pm z_{1 - \alpha/2} \sqrt{\widehat{V}_{jj}} using the
sandwich (or Cox model-based) variance returned by dina_fit().
Usage
## S3 method for class 'dina'
confint(object, parm, level = 0.95, ...)
Arguments
object |
an object of class |
parm |
optional character or integer vector selecting which coefficients to return intervals for; defaults to all. |
level |
confidence level (default |
... |
unused. |
Value
a matrix with two columns giving lower and upper limits, named
in standard confint() style (e.g. "2.5 %", "97.5 %").
Resampling approximation to the splitting consistency rate
Description
Approximates the forestsearch consistency rate for a candidate subgroup
without repeated sample splits and refits. The subgroup is fit once; the
random-split pair of half-sample treatment effects is represented as
beta_hat +/- D, a multiplier sum of the per-subject treatment dfbeta
contributions (see the file header). Handles the Cox (survival) path and the
GLM paths (binary OR/RR/RD, continuous MD, count IRR).
Usage
consistency_resample(
df,
hr.consistency = 1,
consistency_threshold = NULL,
comparison_threshold = NULL,
outcome_type = c("survival", "binary", "continuous", "count"),
effect_measure = NULL,
method = c("both", "closed", "mc"),
multiplier = c("rademacher", "normal", "poisson"),
draws = 1000L,
tte.name = "Y",
event.name = "Event",
treat.name = "Treat",
outcome.name = "Y",
offset.name = NULL,
cox_init = 0,
adjust_covariates = NULL,
adverse_outcome = TRUE,
seed = NULL
)
Arguments
df |
data.frame/data.table with the relevant outcome columns. |
hr.consistency |
Numeric; survival HR consistency threshold (also the
default GLM threshold when |
consistency_threshold |
Numeric or |
comparison_threshold |
Numeric or |
outcome_type |
Character; |
effect_measure |
Character or |
method |
Character; |
multiplier |
Character; |
draws |
Integer; multiplier draws for the Monte Carlo path. |
tte.name, event.name, treat.name |
Character; survival/treatment column names. |
outcome.name, offset.name |
Character; GLM outcome and (rate) offset column names. |
cox_init |
Numeric; Cox warm-start. |
adjust_covariates |
Character vector or |
adverse_outcome |
Logical; GLM adverse-outcome convention (see
|
seed |
Integer or |
Details
multiplier = "rademacher" (default) reproduces the 50/50 Bernoulli split of
run_single_consistency_split(); "normal" is the Lin (1997) Gaussian
multiplier and "poisson" a centred unit-Poisson (DBP) multiplier.
Value
List with rate_closed, rate_mc, beta_hat, sigma_D, n, d,
delta, measure; all-NA rates (with available summaries) if the fit
fails or the configuration is unsupported.
See Also
run_single_consistency_split(), consistency_resample_compare()
Examples
## Not run:
# survival (Cox)
consistency_resample(sub, hr.consistency = 1.0)
# binary odds ratio
consistency_resample(sub, outcome_type = "binary", effect_measure = "OR",
consistency_threshold = 1.0,
outcome.name = "Y", treat.name = "Treat")
# count incidence-rate ratio
consistency_resample(sub, outcome_type = "count", effect_measure = "IRR",
consistency_threshold = 1.0,
outcome.name = "events", offset.name = "fu_time")
## End(Not run)
Validate the consistency approximation against literal splitting (Cox)
Description
Calls the package's run_single_consistency_split() R_true times and
compares the Monte Carlo rate to the closed-form and multiplier
approximations from consistency_resample(). Survival path only; for the
GLM paths use the standalone GLM validator.
Usage
consistency_resample_compare(
df,
hr.consistency = 1,
R_true = 400L,
draws = 1000L,
multiplier = "rademacher",
tte.name = "Y",
event.name = "Event",
treat.name = "Treat",
cox_init = 0,
adjust_covariates = NULL,
seed = NULL
)
Arguments
df |
data.frame/data.table with the relevant outcome columns. |
hr.consistency |
Numeric; survival HR consistency threshold (also the
default GLM threshold when |
R_true |
Integer; number of literal random splits used for the truth. |
draws |
Integer; multiplier draws for the Monte Carlo path. |
multiplier |
Character; |
tte.name, event.name, treat.name |
Character; survival/treatment column names. |
cox_init |
Numeric; Cox warm-start. |
adjust_covariates |
Character vector or |
seed |
Integer or |
Value
One-row data.frame with n, d, beta_hat, sigma_D,
rate_true, valid_true, rate_closed, rate_mc, err_closed,
err_mc.
See Also
consistency_resample(), run_single_consistency_split()
Count ID Occurrences in Bootstrap Sample
Description
Counts the number of times an ID appears in a bootstrap sample.
Usage
count_boot_id(x, dfb)
Arguments
x |
ID value. |
dfb |
Data frame of bootstrap sample. |
Value
Integer count of occurrences.
Examples
df_boot <- data.frame(id = c(1, 2, 1, 3, 1), id_boot = 1:5)
count_boot_id(1, df_boot) # returns 3
count_boot_id(4, df_boot) # returns 0
Comprehensive Wrapper for Cox Spline Analysis with AHR and CDE Plotting
Description
This wrapper function combines Cox spline fitting with comprehensive visualization of Average Hazard Ratios (AHRs) and Controlled Direct Effects (CDEs) as described in the MRCT subgroups analysis documentation.
Usage
cox_ahr_cde_analysis(
df,
tte_name = "os_time",
event_name = "os_event",
treat_name = "treat",
z_name = "biomarker",
loghr_po_name = "loghr_po",
theta1_name = "theta_1",
theta0_name = "theta_0",
spline_df = 3,
alpha = 0.2,
hr_threshold = 0.7,
plot_style = c("combined", "separate", "grid"),
plot_select = c("all", "profile_ahr", "ahr_only"),
save_plots = FALSE,
output_dir = "plots",
verbose = TRUE
)
Arguments
df |
Data frame containing survival data with potential outcomes. |
tte_name |
Character string specifying time-to-event variable name.
Default: |
event_name |
Character string specifying event indicator variable
name. Default: |
treat_name |
Character string specifying treatment variable name.
Default: |
z_name |
Character string specifying continuous covariate/biomarker
name. Default: |
loghr_po_name |
Character string specifying potential outcome log HR
variable. Default: |
theta1_name |
Optional: variable name for theta_1 (treated potential
outcome). Default: |
theta0_name |
Optional: variable name for theta_0 (control potential
outcome). Default: |
spline_df |
Integer degrees of freedom for spline fitting. Default: 3. |
alpha |
Numeric significance level for confidence intervals. Default: 0.20. |
hr_threshold |
Numeric hazard ratio threshold for subgroup
identification, or |
plot_style |
Character: |
plot_select |
Character controlling which panels to display:
|
save_plots |
Logical whether to save plots to file. Default: FALSE. |
output_dir |
Character directory for saving plots. Default:
|
verbose |
Logical for diagnostic output. Default: TRUE. |
Value
List of class "cox_ahr_cde" containing:
- cox_fit
Results from
cox_cs_fitfunction.- ahr_results
AHR calculations for different subgroup definitions.
- cde_results
CDE calculations if theta variables available.
- optimal_cutpoint
Optimal biomarker cutpoint, or
NULLwhenhr_thresholdisNULL.- subgroup_stats
Statistics for recommended and questionable subgroups, or overall-only when
hr_thresholdisNULL.- data
List with z_values, loghr_po, and subgroup assignments.
Examples
## Not run:
# With threshold - full subgroup analysis
results <- cox_ahr_cde_analysis(
df = df_large, z_name = "z_bm",
hr_threshold = 1.25, plot_style = "grid"
)
# Without threshold - pure AHR/CDE curves
results <- cox_ahr_cde_analysis(
df = df_large, z_name = "z_bm",
hr_threshold = NULL, plot_style = "grid"
)
# Compact two-panel without threshold
results <- cox_ahr_cde_analysis(
df = df_large, z_name = "z_bm",
hr_threshold = NULL,
plot_select = "profile_ahr"
)
# Single AHR panel only
results <- cox_ahr_cde_analysis(
df = df_large, z_name = "z_bm",
hr_threshold = 1.25,
plot_select = "ahr_only"
)
## End(Not run)
Fit Cox Model with Cubic Spline for Treatment Effect Heterogeneity
Description
Estimates treatment effects as a function of a continuous covariate using a Cox proportional hazards model with natural cubic splines. The function models treatment-by-covariate interactions to detect effect modification.
Usage
cox_cs_fit(
df,
tte_name = "os_time",
event_name = "os_event",
treat_name = "treat",
strata_name = NULL,
z_name = "bm",
alpha = 0.2,
spline_df = 3,
z_max = Inf,
z_by = 1,
z_window = 0,
z_quantile = 0.9,
show_plot = TRUE,
plot_params = NULL,
truebeta_name = NULL,
verbose = TRUE
)
Arguments
df |
Data frame containing survival data |
tte_name |
Character string specifying time-to-event variable name. Default: "os_time" |
event_name |
Character string specifying event indicator variable name (1=event, 0=censored). Default: "os_event" |
treat_name |
Character string specifying treatment variable name (1=treated, 0=control). Default: "treat" |
strata_name |
Character string specifying stratification variable name. If NULL, no stratification is used. Default: NULL |
z_name |
Character string specifying continuous covariate name for effect modification. Default: "bm" |
alpha |
Numeric value for confidence level (two-sided). Default: 0.20 (80% confidence intervals) |
spline_df |
Integer specifying degrees of freedom for natural spline. Default: 3 |
z_max |
Numeric maximum value for z in predictions. Values beyond this are truncated. Default: Inf (no truncation) |
z_by |
Numeric increment for z values in prediction grid. Default: 1 |
z_window |
Numeric half-width for counting observations near each z value. Default: 0.0 (exact matches only) |
z_quantile |
Numeric quantile (0-1) for upper limit of z profile. Default: 0.90 (90th percentile) |
show_plot |
Logical indicating whether to display plot. Default: TRUE |
plot_params |
List of plotting parameters (see Details). Default: NULL |
truebeta_name |
Character string specifying variable containing true log(HR) values for validation/simulation. Default: NULL |
verbose |
Logical indicating whether to print diagnostic information. Default: TRUE |
Details
Model Structure
The function fits:
h(t|Z,A) = h_0(t) \exp(\beta_0 A + f(Z) + g(Z) \cdot A)
Where:
A is treatment (0/1)
Z is the continuous effect modifier
f(Z) is modeled with natural splines (main effect)
g(Z) is modeled with natural splines (interaction)
The log hazard ratio is:
\beta(Z) = \beta_0 + g(Z)
Plot Parameters
The plot_params argument accepts a list with:
-
xlab: x-axis label -
main_title: plot title -
ylimit: y-axis limits c(min, max) -
y_pad_zero: padding below zero line -
y_delta: extra space for count labels -
cex_legend: legend text size -
cex_count: count text size -
show_cox_primary: show standard Cox estimate line -
show_null: show null effect line (log(HR)=0) -
show_target: show target effect line (e.g., log(0.80))
Value
List containing:
- z_profile
Vector of z values where treatment effect is estimated
- loghr_est
Point estimates of log(HR) at each z value
- loghr_lower
Lower confidence bound
- loghr_upper
Upper confidence bound
- se_loghr
Standard errors of log(HR) estimates
- counts_profile
Number of observations near each z value
- cox_primary
Log(HR) from standard Cox model (no interaction)
- model_fit
The fitted coxph model object
- spline_basis
The natural spline basis object
Examples
## Not run:
# Simulate data
set.seed(123)
df <- data.frame(
os_time = rexp(500, 0.01),
os_event = rbinom(500, 1, 0.7),
treat = rbinom(500, 1, 0.5),
bm = rnorm(500, 50, 10)
)
# Fit model
result <- cox_cs_fit(df, z_name = "bm", alpha = 0.20)
# Custom plotting
result <- cox_cs_fit(
df,
z_name = "bm",
plot_params = list(
xlab = "Biomarker Level",
main_title = "Treatment Effect by Biomarker",
cex_legend = 1.2
)
)
## End(Not run)
Cox model summary for subgroup (OPTIMIZED)
Description
Called in analyze_subgroup() <– SG_tab_estimates
Usage
cox_summary(
Y,
E,
Treat,
Strata = NULL,
use_strata = !is.null(Strata),
return_format = c("formatted", "numeric")
)
Arguments
Y |
Numeric vector of outcome. |
E |
Numeric vector of event indicators. |
Treat |
Numeric vector of treatment indicators. |
Strata |
Vector of strata (optional). |
use_strata |
Logical. Whether to use strata in the model (default: TRUE if Strata provided). |
return_format |
Character. "formatted" (default) or "numeric" for downstream use. |
Details
Calculates hazard ratio and confidence interval for a subgroup using Cox regression. Optimized version with reduced overhead and better error handling.
Value
Character string with formatted HR and CI (or numeric vector if return_format="numeric").
Examples
## Not run:
library(survival)
cox_summary(
Y = gbsg$rfstime / 30.4375,
E = gbsg$status,
Treat = gbsg$hormon
)
## End(Not run)
Batch Cox summaries with caching
Description
For repeated calls with the same data structure but different subsets, this version pre-processes the data structure once.
Usage
cox_summary_batch(
Y,
E,
Treat,
Strata = NULL,
subset_indices,
return_format = c("formatted", "numeric")
)
Arguments
Y |
Numeric vector of outcome (full dataset). |
E |
Numeric vector of event indicators (full dataset). |
Treat |
Numeric vector of treatment indicators (full dataset). |
Strata |
Vector of strata (optional, full dataset). |
subset_indices |
List of integer vectors, each defining a subset to analyze. |
return_format |
Character. "formatted" or "numeric". |
Value
List of results, one per subset.
Cox model summary for subgroup
Description
Called in analyze_subgroup() <– SG_tab_estimates
Usage
cox_summary_legacy(Y, E, Treat, Strata)
Arguments
Y |
Numeric vector of outcome. |
E |
Numeric vector of event indicators. |
Treat |
Numeric vector of treatment indicators. |
Strata |
Vector of strata (optional). |
Details
Calculates hazard ratio and confidence interval for a subgroup using Cox regression.
Value
Character string with formatted HR and CI.
Cox model summary for subgroup - vectorized version
Description
Efficiently processes multiple subgroups at once. Useful when analyzing many subgroups (e.g., in cross-validation).
Usage
cox_summary_vectorized(
data,
outcome_col,
event_col,
treat_col,
strata_col = NULL,
subgroup_col = "subgroup",
return_format = c("formatted", "numeric")
)
Arguments
data |
Data frame with columns for Y, E, Treat, and optionally Strata. |
outcome_col |
Character. Name of outcome column. |
event_col |
Character. Name of event column. |
treat_col |
Character. Name of treatment column. |
strata_col |
Character. Name of strata column (optional). |
subgroup_col |
Character. Name of subgroup indicator column. |
return_format |
Character. "formatted" or "numeric". |
Value
Data frame with one row per subgroup and HR results.
Calculate Bootstrap Table Caption
Description
Generates an interpretive caption for bootstrap results table.
Usage
create_bootstrap_caption(est.scale, nb_boots, boot_success_rate)
Arguments
est.scale |
Character. "hr" or "1/hr" |
nb_boots |
Integer. Number of bootstrap iterations |
boot_success_rate |
Numeric. Proportion successful |
Value
Character string with caption
Create Bootstrap Diagnostic Plots
Description
Generates diagnostic visualization plots for bootstrap analysis.
Usage
create_bootstrap_diagnostic_plots(
results,
H_estimates,
Hc_estimates,
overall_timing = NULL,
effect_label = "HR"
)
Arguments
results |
Data frame with bootstrap results |
H_estimates |
List with H subgroup estimates |
Hc_estimates |
List with Hc subgroup estimates |
overall_timing |
List with overall timing information (optional) |
Value
List of ggplot2 objects
Create Data Generating Mechanism for MRCT Simulations
Description
Wrapper function to create a data generating mechanism (DGM) for MRCT
simulation scenarios using generate_aft_dgm_flex.
Usage
create_dgm_for_mrct(
df_case,
model_type = c("alt", "null"),
log_hrs = NULL,
confounder_var = NULL,
confounder_effect = NULL,
include_regA = TRUE,
verbose = FALSE
)
Arguments
df_case |
Data frame containing case study data |
model_type |
Character. Either "alt" (alternative hypothesis with heterogeneous treatment effects) or "null" (uniform treatment effect) |
log_hrs |
Numeric vector. Log hazard ratios for spline specification. If NULL, defaults are used based on model_type |
confounder_var |
Character. Name of a confounder variable to include with a forced prognostic effect. Default: NULL (no forced effect) |
confounder_effect |
Numeric. Log hazard ratio for confounder_var effect. Only used if confounder_var is specified |
include_regA |
Logical. Include regA as a factor in the model. Default: TRUE |
verbose |
Logical. Print detailed output. Default: FALSE |
Details
Model Types
- alt
Alternative hypothesis: Treatment effect varies by biomarker level (heterogeneous treatment effect). Default log_hrs create HR ranging from 2.0 (harm) to 0.5 (benefit) across biomarker range
- null
Null hypothesis: Uniform treatment effect regardless of biomarker level. Default log_hrs = log(0.7) uniformly
Confounder Effects
By default, NO prognostic confounder effect is forced. The confounder_var and confounder_effect parameters allow optionally specifying ANY baseline covariate to have a fixed prognostic effect in the outcome model.
The regA variable (region indicator) is included as a factor by default but without a forced effect - its coefficient is estimated from data.
Value
An object of class "aft_dgm_flex" for use with
simulate_from_dgm and mrct_region_sims
See Also
generate_aft_dgm_flex for underlying DGM creation
mrct_region_sims for running simulations with the DGM
Examples
## Not run:
# Prepare data
df_case <- read.csv("data/dfsynthetic.csv")
df_case$regA <- df_case$region_asia
# Alternative hypothesis (heterogeneous treatment effect)
dgm_alt <- create_dgm_for_mrct(
df_case = df_case,
model_type = "alt",
log_hrs = log(c(3, 1.25, 0.50)),
verbose = TRUE
)
# Null hypothesis (uniform effect)
dgm_null <- create_dgm_for_mrct(
df_case = df_case,
model_type = "null",
verbose = TRUE
)
# With forced confounder effect
dgm_conf <- create_dgm_for_mrct(
df_case = df_case,
model_type = "alt",
confounder_var = "prior_treat",
confounder_effect = log(1.5),
verbose = TRUE
)
## End(Not run)
Create Factor Summary Tables from Bootstrap Results
Description
Generates formatted GT tables summarizing factor frequencies from bootstrap subgroup analysis. Creates two complementary tables: one showing factor selection frequencies within each position (M.1, M.2, etc.), and another showing overall factor frequencies across all positions.
Usage
create_factor_summary_tables(factor_freq, n_found, min_percent = 2)
Arguments
factor_freq |
Data.frame or data.table. Factor frequency table from
|
n_found |
Integer. Number of successful bootstrap iterations (where a subgroup was identified). Used to calculate overall percentages. |
min_percent |
Numeric. Minimum percentage threshold for including factors in the tables. Factors with selection frequencies below this threshold are excluded. Default is 2 (i.e., 2%). |
Value
A list with up to two GT table objects:
by_positionGT table showing factor frequencies within each position. Percentages represent conditional probability of factor selection given that the position was populated. Within each position, percentages sum to approximately 100% (may not sum exactly to 100% after filtering).
overallGT table showing total factor frequencies across all positions. Includes additional columns indicating which positions each factor appeared in and how many unique positions used the factor. Percentages represent proportion of successful iterations where the factor appeared in any position.
If no factors meet the minimum threshold, the corresponding table element will be NULL.
Note
This function requires the gt package for table creation. The overall table also requires dplyr for data aggregation. If dplyr is not available, only the position-specific table will be created and the overall element will be NULL.
Always check for NULL before using the returned tables:
if (!is.null(factor_tables$by_position)) {
print(factor_tables$by_position)
}
If all factors have percentages below min_percent, both table elements
will be NULL.
See Also
-
summarize_bootstrap_subgroupsfor generating the factor_freq input -
format_subgroup_summary_tablesfor creating all subgroup summary tables -
summarize_bootstrap_resultsfor complete bootstrap analysis workflow -
forestsearch_bootstrap_dofuturefor running bootstrap analysis
Create Forest Plot Theme with Size Controls
Description
Creates a forestploter theme with parameters that control overall plot sizing and appearance. This is the primary way to control how large the forest plot renders.
Usage
create_forest_theme(
base_size = 10,
scale = 1,
row_padding = NULL,
ci_pch = 15,
ci_lwd = NULL,
ci_Theight = NULL,
ci_col = "black",
header_fontsize = NULL,
body_fontsize = NULL,
footnote_fontsize = NULL,
footnote_col = "darkcyan",
title_fontsize = NULL,
cv_fontsize = NULL,
cv_col = "gray30",
refline_lwd = NULL,
refline_lty = "dashed",
refline_col = "gray30",
vertline_lwd = NULL,
vertline_lty = "dashed",
vertline_col = "gray20",
arrow_type = "closed",
arrow_col = "black",
summary_fill = "black",
summary_col = "black"
)
Arguments
base_size |
Numeric. Base font size in points. This is the primary scaling parameter - increasing it will proportionally scale all fonts, row padding, and line widths. Default: 10. |
scale |
Numeric. Additional scaling multiplier applied on top of base_size. Use for quick overall scaling. Default: 1.0. |
row_padding |
Numeric vector of length 2. Padding around row content in mm as c(vertical, horizontal). If NULL, auto-calculated from base_size. Default: NULL. |
ci_pch |
Integer. Point character for CI. 15=square, 16=circle, 18=diamond. Default: 15. |
ci_lwd |
Numeric. Line width for CI lines. If NULL, auto-calculated from base_size. Default: NULL. |
ci_Theight |
Numeric. Height of T-bar ends on CI. If NULL, auto-calculated from base_size. Default: NULL. |
ci_col |
Character. Color for CI lines and points. Default: "black". |
header_fontsize |
Numeric. Font size for column headers. If NULL, auto-calculated as base_size * scale + 1. Default: NULL. |
body_fontsize |
Numeric. Font size for body text. If NULL, auto-calculated as base_size * scale. Default: NULL. |
footnote_fontsize |
Numeric. Font size for footnotes. If NULL, auto-calculated as base_size * scale - 1. Default: NULL. |
footnote_col |
Character. Color for footnote text. Default: "darkcyan". |
title_fontsize |
Numeric. Font size for title. If NULL, auto-calculated as base_size * scale + 4. Default: NULL. |
cv_fontsize |
Numeric. Font size for CV annotation text. If NULL, auto-calculated as base_size * scale. Default: NULL. |
cv_col |
Character. Color for CV annotation text. Default: "gray30". |
refline_lwd |
Numeric. Reference line width. If NULL, auto-calculated. Default: NULL. |
refline_lty |
Character. Reference line type. Default: "dashed". |
refline_col |
Character. Reference line color. Default: "gray30". |
vertline_lwd |
Numeric. Vertical line width. If NULL, auto-calculated. Default: NULL. |
vertline_lty |
Character. Vertical line type. Default: "dashed". |
vertline_col |
Character. Vertical line color. Default: "gray20". |
arrow_type |
Character. Arrow type: "open" or "closed". Default: "closed". |
arrow_col |
Character. Arrow color. Default: "black". |
summary_fill |
Character. Fill color for summary diamonds. Default: "black". |
summary_col |
Character. Border color for summary diamonds. Default: "black". |
Details
The base_size parameter is the primary way to control plot size.
When you change base_size, the following are automatically scaled:
All font sizes (body, header, footnote, CV, title)
Row padding (vertical and horizontal)
CI line width and T-bar height
Reference and vertical line widths
The scaling formula uses base_size = 10 as the reference point:
base_size = 10: Default sizing
base_size = 12: 20% larger
base_size = 14: 40% larger
base_size = 16: 60% larger
You can override any individual parameter by specifying it explicitly.
The theme does NOT set row background colors - those are determined
automatically by plot_subgroup_results_forestplot() based on
row types (ITT, reference, posthoc, etc.).
Value
A list of class "fs_forest_theme" containing all theme parameters.
See Also
plot_subgroup_results_forestplot, render_forestplot
Examples
## Not run:
# Simple: just increase base_size for larger plot
large_theme <- create_forest_theme(base_size = 14)
# Or use scale for quick adjustment
large_theme <- create_forest_theme(base_size = 10, scale = 1.4)
# Fine-tune specific elements
custom_theme <- create_forest_theme(
base_size = 14,
cv_fontsize = 12, # Override auto-calculated CV font size
ci_lwd = 2.5 # Override auto-calculated CI line width
)
# Use with plot_subgroup_results_forestplot
result <- plot_subgroup_results_forestplot(
fs_results = list(fs.est = fs, fs_bc = fs_bc),
df_analysis = df.analysis,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
theme = large_theme
)
render_forestplot(result)
## End(Not run)
Create Subgroup Indicator Columns from ForestSearch
Description
Internal helper to create Qrecommend and Brecommend indicator columns.
Usage
create_fs_subgroup_indicators(
df,
fs.est,
col_names = c("Qrecommend", "Brecommend"),
verbose = FALSE
)
Arguments
df |
Data frame to modify. |
fs.est |
A forestsearch object. |
col_names |
Character vector of length 2. Names for the indicator columns: first for harm/questionable (treat.recommend == 0), second for benefit/recommend (treat.recommend == 1). Default: c("Qrecommend", "Brecommend") |
verbose |
Logical. Print diagnostic messages. |
Value
Modified data frame with indicator columns.
Create GBSG-Based AFT Data Generating Mechanism
Description
Creates a data generating mechanism (DGM) for survival simulations based on the German Breast Cancer Study Group (GBSG) dataset. Supports heterogeneous treatment effects via treatment-subgroup interactions.
Usage
create_gbsg_dgm(
model = c("alt", "null"),
k_treat = 1,
k_inter = 1,
k_z3 = 1,
z1_quantile = 0.25,
n_super = DEFAULT_N_SUPER,
cens_type = c("weibull", "uniform"),
use_rand_params = FALSE,
seed = SEED_BASE,
verbose = FALSE
)
Arguments
model |
Character. Either "alt" for alternative hypothesis with heterogeneous treatment effects, or "null" for uniform treatment effect. Default: "alt" |
k_treat |
Numeric. Treatment effect multiplier applied to the treatment coefficient from the fitted AFT model. Values > 1 strengthen the treatment effect. Default: 1 |
k_inter |
Numeric. Interaction effect multiplier for the treatment-subgroup interaction (z1 * z3). Only used when model = "alt". Higher values create more heterogeneity between HR(H) and HR(Hc). Default: 1 |
k_z3 |
Numeric. Effect multiplier for the z3 (menopausal status) coefficient. Default: 1 |
z1_quantile |
Numeric. Quantile threshold for z1 (estrogen receptor). Observations with ER <= quantile are coded as z1 = 1. Default: 0.25 |
n_super |
Integer. Size of super-population for empirical HR estimation. Default: 5000 |
cens_type |
Character. Censoring distribution type: "weibull" or "uniform". Default: "weibull" |
use_rand_params |
Logical. If TRUE, modifies confounder coefficients using estimates from randomized subset (meno == 0). Default: FALSE |
seed |
Integer. Random seed for super-population generation. Default: 8316951 |
verbose |
Logical. Print diagnostic information. Default: FALSE |
Details
This version is aligned with generate_aft_dgm_flex() and
calculate_hazard_ratios() methodology, computing individual-level
potential outcomes and average hazard ratios (AHR).
Subgroup Definition
The harm subgroup H is defined as: z1 = 1 AND z3 = 1, where:
z1: Low estrogen receptor (ER <= 25th percentile by default)
z3: Premenopausal status (meno == 0)
Model Specification
The AFT model uses covariates: treat, z1, z2, z3, z4, z5, and (for "alt") the interaction zh = treat * z1 * z3.
Interaction Effect (k_inter)
The k_inter parameter modifies the zh coefficient in the AFT model:
gamma[zh] <- k_inter * gamma[zh]
This affects the hazard ratio for the harm subgroup:
HR(H) = exp(-gamma[treat]/sigma - gamma[zh]/sigma)
HR(Hc) = exp(-gamma[treat]/sigma)
When k_inter = 0, HR(H) = HR(Hc) (no heterogeneity).
Alignment with generate_aft_dgm_flex
This function now computes:
theta_0: Log-hazard contribution under control
theta_1: Log-hazard contribution under treatment
loghr_po: Individual causal log hazard ratio (theta_1 - theta_0)
AHR metrics: exp(mean(loghr_po)) for overall and subgroups
Value
A list of class "gbsg_dgm" containing:
- df_super_rand
Data frame with randomized super-population including potential outcomes (theta_0, theta_1, loghr_po)
- hr_H_true
Empirical hazard ratio in harm subgroup (Cox-based)
- hr_Hc_true
Empirical hazard ratio in complement subgroup (Cox-based)
- hr_causal
Overall causal (ITT) hazard ratio (Cox-based)
- AHR
Overall average hazard ratio (from loghr_po)
- AHR_H_true
Average hazard ratio in harm subgroup
- AHR_Hc_true
Average hazard ratio in complement subgroup
- hazard_ratios
List matching generate_aft_dgm_flex output format
- model_params
List with AFT model parameters (mu, sigma, gamma, etc.)
- cens_params
List with censoring model parameters
- subgroup_info
List with subgroup definitions and true factor names
- analysis_vars
Character vector of analysis variable names
- model_type
Character indicating "alt" or "null"
See Also
simulate_from_gbsg_dgm for generating data from the DGM
calibrate_k_inter for finding k_inter to achieve target HR
Examples
## Not run:
# Alternative hypothesis with default parameters
dgm_alt <- create_gbsg_dgm(model = "alt", verbose = TRUE)
# Null hypothesis
dgm_null <- create_gbsg_dgm(model = "null", verbose = TRUE)
# Custom subgroup HR via k_inter
dgm_custom <- create_gbsg_dgm(
model = "alt",
k_treat = 1.2,
k_inter = 2.0,
verbose = TRUE
)
# Access AHR metrics (aligned with generate_aft_dgm_flex)
dgm_alt$hazard_ratios$AHR_harm
dgm_alt$hazard_ratios$AHR_no_harm
## End(Not run)
Compute GLM Effect Estimate for a Single Subgroup
Description
Fits a GLM within the specified subgroup and returns a one-row data frame containing the effect estimate, confidence interval, sample size, and event count. Supports binary, continuous, and count (with offset) outcomes.
Usage
create_glm_row(
df,
outcome.name,
treat.name = "treat",
effect_measure = "log_OR",
offset.name = NULL,
overdispersion = "none",
conf.level = 0.95,
min_arm_n = 5L,
min_arm_events = 3L,
verbose = FALSE
)
Arguments
df |
Data frame for the subgroup (already subset). |
outcome.name |
Character. Name of the outcome column. |
treat.name |
Character. Name of the treatment indicator. |
effect_measure |
Character. Effect measure; see
|
offset.name |
Character or |
overdispersion |
Character. One of |
conf.level |
Numeric. Confidence level. Default: |
min_arm_n |
Integer. Minimum observations per arm; returns |
min_arm_events |
Integer. For binary/count outcomes, minimum events
per arm; returns |
verbose |
Logical. Default: |
Value
A one-row data frame with columns:
N, n_treat, n_control,
events_treat, events_control (for binary/count),
est, lower, upper, se, effect_measure.
See Also
Helper Functions for GRF Subgroup Analysis
Description
This file contains helper functions used by grf.subg.harm.survival() to improve readability and modularity. Create GRF configuration object
Usage
create_grf_config(
frac.tau,
n.min,
dmin.grf,
RCT,
sg.criterion,
maxdepth,
seedit
)
Arguments
frac.tau |
Numeric. Fraction of tau for GRF horizon |
n.min |
Integer. Minimum subgroup size |
dmin.grf |
Numeric. Minimum difference in subgroup mean |
RCT |
Logical. Is the data from a randomized controlled trial? |
sg.criterion |
Character. Subgroup selection criterion |
maxdepth |
Integer. Maximum tree depth |
seedit |
Integer. Random seed |
Details
Creates a configuration object to organize GRF parameters
Value
List with configuration parameters
Examples
cfg <- create_grf_config(frac.tau = 0.6, n.min = 60, dmin.grf = 6,
RCT = TRUE, sg.criterion = "mDiff",
maxdepth = 2, seedit = 42L)
str(cfg)
Create result object when no subgroup is found
Description
Builds result object for cases where no valid subgroup is identified
Usage
create_null_result(data, values, trees, config)
Arguments
data |
Data frame. Original data |
values |
Data frame. Node metrics (may be empty) |
trees |
List. Fitted policy trees |
config |
List. GRF configuration |
Value
List with limited GRF results
Create Reference Subgroup Indicator Columns
Description
Creates indicator columns (0/1) in the data frame for each reference subgroup based on the provided subset expressions.
Usage
create_reference_subgroup_columns(df, ref_subgroups, verbose = FALSE)
Arguments
df |
Data frame to modify. |
ref_subgroups |
Named list of reference subgroup definitions.
Each element should have |
verbose |
Logical. Print diagnostic messages. |
Value
List with modified df, cols, labels, and colors vectors.
Create Result Row
Description
Create Result Row
Usage
create_result_row(kk, covs.in, nx, event_counts, cox_result)
Create Sample Size Table for Multiple Scenarios
Description
Generates a table of required sample sizes for different combinations of true hazard ratios and censoring proportions.
Usage
create_sample_size_table(
theta_values,
prop_cens_values,
target_power = 0.8,
hr_threshold = 1.25,
verbose = TRUE
)
Arguments
theta_values |
Numeric vector. True hazard ratios to evaluate. |
prop_cens_values |
Numeric vector. Censoring proportions to evaluate. |
target_power |
Numeric. Target detection probability. Default: 0.80 |
hr_threshold |
Numeric. HR threshold. Default: 1.25 |
verbose |
Logical. Print progress. Default: TRUE |
Value
A data.frame with columns: theta, prop_cens, n_required, achieved_power
Examples
## Not run:
ss_table <- create_sample_size_table(
theta_values = c(1.5, 1.75, 2.0, 2.5),
prop_cens_values = c(0.2, 0.3, 0.4),
target_power = 0.80
)
print(ss_table)
## End(Not run)
Create Spline Variables
Description
Create Spline Variables
Usage
create_spline_variables(df_work, spline_var, knot)
Create Subgroup Indicator from Factor Definitions
Description
Parses factor definitions (e.g., "v1.1", "grade3.1") and creates a binary indicator for subgroup membership.
Usage
create_subgroup_indicator(df, sg_factors)
Arguments
df |
Data frame containing the variables |
sg_factors |
Character vector of factor definitions |
Value
Integer vector (1 = in subgroup, 0 = not in subgroup)
Create Subgroup Summary Data Frame for Forest Plot
Description
Creates a data frame suitable for forestploter from multiple subgroup analyses. This is a more flexible alternative for complex subgroup configurations.
Usage
create_subgroup_summary_df(
df_analysis,
subgroups,
outcome.name,
event.name,
treat.name,
E.name = "E",
C.name = "C",
fs_bc_list = NULL,
fs_kfold_list = NULL,
conf.level = 0.95
)
Arguments
df_analysis |
Data frame. The analysis dataset. |
subgroups |
Named list of subgroup definitions. |
outcome.name |
Character. Name of survival time variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
E.name |
Character. Label for experimental arm. |
C.name |
Character. Label for control arm. |
fs_bc_list |
List. Named list of bootstrap results for each subgroup. |
fs_kfold_list |
List. Named list of k-fold results for each subgroup. |
conf.level |
Numeric. Confidence level for intervals (default: 0.95). |
Value
Data frame with HR estimates for all subgroups.
Create result object for successful subgroup identification
Description
Builds comprehensive result object when a subgroup is found
Usage
create_success_result(
data,
best_subgroup,
trees,
tree_cuts,
selected_tree,
sg_harm_id,
values,
config
)
Arguments
data |
Data frame. Original data with subgroup assignments |
best_subgroup |
Data frame row. Selected subgroup information |
trees |
List. All fitted policy trees |
tree_cuts |
List. Cut information from trees |
selected_tree |
Policy tree. The tree that identified the subgroup |
sg_harm_id |
Character. Expression defining the subgroup |
values |
Data frame. All node metrics |
config |
List. GRF configuration |
Value
List with complete GRF results
Create Enhanced Summary Table for Baseline Characteristics
Description
Generates a formatted summary table comparing baseline characteristics between treatment arms. Supports continuous, categorical, and binary variables with p-values, standardized mean differences (SMD), and missing data summaries.
Usage
create_summary_table(
data,
treat_var = "treat",
vars_continuous = NULL,
vars_categorical = NULL,
vars_binary = NULL,
var_labels = NULL,
digits = 1,
show_pvalue = TRUE,
show_smd = TRUE,
show_missing = TRUE,
table_title = "Baseline Characteristics by Treatment Arm",
table_subtitle = NULL,
source_note = NULL,
font_size = 12,
header_font_size = 14,
footnote_font_size = 10,
use_alternating_rows = TRUE,
stripe_color = "#f9f9f9",
indent_size = 20,
highlight_pval = 0.05,
highlight_smd = 0.2,
highlight_color = "#fff3cd",
compact_mode = FALSE,
column_width_var = 200,
column_width_stats = 120,
show_column_borders = FALSE,
custom_css = NULL
)
Arguments
data |
Data frame containing the analysis data |
treat_var |
Character. Name of treatment variable (must have 2 levels) |
vars_continuous |
Character vector. Names of continuous variables |
vars_categorical |
Character vector. Names of categorical variables |
vars_binary |
Character vector. Names of binary (0/1) variables |
var_labels |
Named list. Custom labels for variables (optional) |
digits |
Integer. Number of decimal places for continuous variables |
show_pvalue |
Logical. Include p-values column |
show_smd |
Logical. Include SMD (effect size) column |
show_missing |
Logical. Include missing data rows |
table_title |
Character. Main title for the table |
table_subtitle |
Character. Subtitle for the table (optional) |
source_note |
Character. Source note at bottom (optional) |
font_size |
Numeric. Base font size in pixels (default: 12) |
header_font_size |
Numeric. Header font size in pixels (default: 14) |
footnote_font_size |
Numeric. Footnote font size in pixels (default: 10) |
use_alternating_rows |
Logical. Apply zebra striping (default: TRUE) |
stripe_color |
Character. Color for alternating rows (default: "#f9f9f9") |
indent_size |
Numeric. Indentation for sub-levels in pixels (default: 20) |
highlight_pval |
Numeric. Highlight p-values below this threshold (default: 0.05) |
highlight_smd |
Numeric. Highlight SMD values above this threshold (default: 0.2) |
highlight_color |
Character. Color for highlighting (default: "#fff3cd") |
compact_mode |
Logical. Reduce spacing for compact display (default: FALSE) |
column_width_var |
Numeric. Width for Variable column in pixels (default: 200) |
column_width_stats |
Numeric. Width for stat columns in pixels (default: 120) |
show_column_borders |
Logical. Show vertical column borders (default: FALSE) |
custom_css |
Character. Additional custom CSS styling (optional) |
Details
Binary variables specified via vars_binary display a single row
showing the count and proportion for the "1" level. Categorical variables
specified via vars_categorical that happen to be binary-coded (i.e.,
have exactly two levels: 0 and 1) are automatically detected and displayed
in the same compact single-row format, showing only the "1" proportion.
Value
A gt table object (or data frame if gt not available)
Examples
## Not run:
# Basic usage
create_summary_table(
data = trial_data,
treat_var = "treatment",
vars_continuous = c("age", "bmi"),
vars_categorical = c("sex", "stage")
)
# Customized appearance
create_summary_table(
data = trial_data,
treat_var = "treatment",
vars_continuous = c("age", "bmi"),
vars_categorical = c("sex", "stage"),
font_size = 11,
header_font_size = 13,
use_alternating_rows = TRUE,
highlight_pval = 0.05,
compact_mode = TRUE
)
## End(Not run)
Preset: Compact Table
Description
Preset: Compact Table
Usage
create_summary_table_compact(...)
Arguments
... |
Arguments passed to create_summary_table() |
Preset: Minimal Table (No Highlighting, No Alternating)
Description
Preset: Minimal Table (No Highlighting, No Alternating)
Usage
create_summary_table_minimal(...)
Arguments
... |
Arguments passed to create_summary_table() |
Preset: Presentation Table (Large Fonts)
Description
Preset: Presentation Table (Large Fonts)
Usage
create_summary_table_presentation(...)
Arguments
... |
Arguments passed to create_summary_table() |
Preset: Publication-Ready Table
Description
Preset: Publication-Ready Table
Usage
create_summary_table_publication(...)
Arguments
... |
Arguments passed to create_summary_table() |
Create Timing Summary Table
Description
Creates a data frame summarizing bootstrap timing information.
Usage
create_timing_summary_table(
overall_timing,
iteration_stats,
fs_stats,
overhead_stats,
nb_boots,
boot_success_rate
)
Arguments
overall_timing |
List. Overall timing statistics |
iteration_stats |
List. Per-iteration timing statistics |
fs_stats |
List. ForestSearch-specific timing statistics |
overhead_stats |
List. Overhead timing statistics |
nb_boots |
Integer. Number of bootstrap iterations |
boot_success_rate |
Numeric. Proportion of successful bootstraps |
Value
Data frame with timing summary
Discretize Continuous Variable into Quantile-Based Categories
Description
Discretize Continuous Variable into Quantile-Based Categories
Usage
cut_numeric(x, probs = c(0.25, 0.5, 0.75))
Arguments
x |
Numeric vector to discretize |
probs |
Numeric vector of probabilities for quantile breaks. Default: c(0.25, 0.5, 0.75) creates quartiles coded as 1, 2, 3, 4 |
Value
Integer vector with category codes (1 = lowest, max = highest)
Discretize Continuous Variable by Size Categories
Description
Discretize Continuous Variable by Size Categories
Usage
cut_size(x, breaks = c(20, 50))
Arguments
x |
Numeric vector (typically tumor size) |
breaks |
Numeric vector of breakpoints. Default: c(20, 50) |
Value
Integer vector with category codes
Generate cut expressions for a variable
Description
For a continuous variable, returns expressions for mean, median, qlow, and qhigh cuts.
Usage
cut_var(x)
Arguments
x |
Character. Variable name. |
Value
Character vector of cut expressions.
Generate J-quantile cut expressions for a continuous variable
Description
For a continuous variable, returns J cut expressions of the
form "X <= qj(X, k, J + 1)" for k = 1, ..., J. These
J cut points are placed at the (k/(J+1))-th empirical
quantiles of X and partition its range into J + 1
non-overlapping intervals
[\min(X), c_1),\ [c_1, c_2),\ \ldots,\ [c_J, \max(X)]
where c_k is the (k/(J+1))-th quantile of X. Note the cut
expressions themselves are nested half-spaces (each is a subset of
the next); ForestSearch consumes them as binary candidate factors
and combines them via intersection during the search.
Usage
cut_var_jq(x, J)
Arguments
x |
Character. Variable name. |
J |
Integer >= 1. Number of binary cut expressions to emit.
The resulting partition has |
Details
Expressions are emitted in deferred form (with literal qj(...)
calls inside the string) so that they are correctly recomputed when
the same expression is processed against a different data subset
(e.g., a bootstrap replicate). They are subsequently resolved to
literal numerics by process_conf_force_expr().
Value
Character vector of J cut expressions.
Compare Multiple CV Results
Description
Creates a comparison table from multiple cross-validation runs with different configurations.
Usage
cv_compare_results(
cv_list,
metrics = c("all", "finding", "agreement"),
show_percentages = TRUE,
digits = 1,
use_gt = TRUE
)
Arguments
cv_list |
Named list of cv_result objects from |
metrics |
Character vector. Which metrics to include. Options: "finding", "agreement", "all". Default: "all". |
show_percentages |
Logical. Display as percentages. Default: TRUE. |
digits |
Integer. Decimal places. Default: 1. |
use_gt |
Logical. Return gt table if TRUE. Default: TRUE. |
Value
A gt table or data.frame comparing CV results across configurations.
Examples
## Not run:
# Compare CV results from different configurations
cv_comparison <- cv_compare_results(
cv_list = list(
"maxk=1" = cv_maxk1,
"maxk=2" = cv_maxk2,
"maxk=3" = cv_maxk3
)
)
## End(Not run)
Create Metrics Tables for Cross-Validation Results
Description
Formats the find_summary and sens_summary outputs from
forestsearch_tenfold or forestsearch_Kfold
into publication-ready gt tables.
Usage
cv_metrics_tables(
cv_result,
sg_definition = NULL,
title = "Cross-Validation Metrics",
show_percentages = TRUE,
digits = 1,
include_raw = FALSE,
table_style = c("combined", "separate", "minimal"),
use_gt = TRUE
)
Arguments
cv_result |
List. Result from |
sg_definition |
Character vector. Subgroup factor definitions for
labeling (optional). If NULL, extracted from |
title |
Character. Main title for combined table. Default: "Cross-Validation Metrics". |
show_percentages |
Logical. Display metrics as percentages (0-100) instead of proportions (0-1). Default: TRUE. |
digits |
Integer. Decimal places for formatting. Default: 1. |
include_raw |
Logical. Include raw matrices ( |
table_style |
Character. One of "combined", "separate", or "minimal".
Default: "combined". |
use_gt |
Logical. Return gt table(s) if TRUE, data.frame(s) if FALSE. Default: TRUE. |
Value
Depending on table_style:
"combined": A single gt table (or data.frame)
"separate": A list with
agreement_tableandfinding_table"minimal": A single-row gt table (or data.frame)
If include_raw = TRUE, also includes sens_out and find_out
matrices in the returned list.
See Also
cv_summary_tables for formatting forestsearch_KfoldOut(outall=TRUE) results
Examples
## Not run:
# After running forestsearch_tenfold
tenfold_results <- forestsearch_tenfold(
fs.est = fs_result,
sims = 100,
Kfolds = 10
)
# Create combined metrics table
cv_tables <- cv_metrics_tables(tenfold_results)
cv_tables
# Create separate tables
cv_tables <- cv_metrics_tables(tenfold_results, table_style = "separate")
cv_tables$agreement_table
cv_tables$finding_table
# Minimal one-row summary
cv_metrics_tables(tenfold_results, table_style = "minimal")
## End(Not run)
Create Summary Tables from forestsearch_KfoldOut Results
Description
Formats the detailed output from forestsearch_KfoldOut(outall=TRUE)
into publication-ready gt tables. This includes ITT estimates, original subgroup
estimates, and K-fold subgroup estimates.
Usage
cv_summary_tables(
kfold_out,
title = "Cross-Validation Summary",
subtitle = NULL,
show_metrics = TRUE,
digits = 3,
font_size = 12,
use_gt = TRUE
)
Arguments
kfold_out |
List. Result from |
title |
Character. Main title for combined table. Default: "Cross-Validation Summary". |
subtitle |
Character. Subtitle for table. Default: NULL (auto-generated). |
show_metrics |
Logical. Include agreement and finding metrics in output. Default: TRUE. |
digits |
Integer. Decimal places for numeric formatting. Default: 3. |
font_size |
Integer. Font size in pixels. Default: 12. |
use_gt |
Logical. Return gt table if TRUE, data.frame if FALSE. Default: TRUE. |
Value
If use_gt = TRUE, returns a list with gt table objects:
-
combined_table: Combined ITT and subgroup estimates -
itt_table: ITT estimates only -
original_table: Original full-data subgroup estimates -
kfold_table: K-fold subgroup estimates -
metrics_table: Agreement and finding metrics (ifshow_metrics = TRUE)
If use_gt = FALSE, returns equivalent data.frames.
See Also
cv_metrics_tables for formatting forestsearch_tenfold() results
Examples
## Not run:
# Run K-fold CV
cv_results <- forestsearch_Kfold(fs.est = fs_result, Kfolds = 10)
# Get detailed output
kfold_out <- forestsearch_KfoldOut(cv_results, outall = TRUE)
# Create summary tables
cv_tables <- cv_summary_tables(kfold_out)
cv_tables$combined_table
cv_tables$metrics_table
## End(Not run)
Create Compact CV Summary Text
Description
Generates a compact text string summarizing CV results, suitable for annotations in plots or reports.
Usage
cv_summary_text(
cv_result,
est.scale = "hr",
include_finding = TRUE,
include_agreement = TRUE
)
Arguments
cv_result |
List. Result from |
est.scale |
Character. "hr" or "1/hr" to determine label orientation. Default: "hr". |
include_finding |
Logical. Include subgroup finding rate. Default: TRUE. |
include_agreement |
Logical |
Value
Character string with formatted CV metrics.
Examples
## Not run:
cv_text <- cv_summary_text(tenfold_results)
# Returns: "CV found = 95%, Agree(+,-) = 88%, 92%"
## End(Not run)
Effective Information for Binary Outcomes (Logistic)
Description
Computes d_{\text{eff}} = n_H \bar{p} (1 - \bar{p}) from the
Fisher information of the logistic regression treatment coefficient.
Usage
d_eff_binary(n_sg, p_event)
Arguments
n_sg |
Integer. Subgroup sample size. |
p_event |
Numeric. Event probability (0–1), averaged across arms. |
Details
For the model \text{logit}(p_i) = \beta_0 + \beta V_i with
balanced arms (n_1 = n_0 = n_H/2):
\text{Var}(\hat\beta) = \frac{4}{n_H \bar{p}(1-\bar{p})}
Value
Numeric. Effective information.
Examples
d_eff_binary(n_sg = 100, p_event = 0.30)
# 21
Effective Information for Continuous Outcomes (Gaussian)
Description
Computes d_{\text{eff}} = n_H / \sigma_Y^2 from the
Fisher information of the Gaussian linear model treatment coefficient.
Usage
d_eff_continuous(n_sg, sigma_y)
Arguments
n_sg |
Integer. Subgroup sample size. |
sigma_y |
Numeric. Residual standard deviation of the outcome. |
Details
For the model Y_i = \beta_0 + \beta V_i + \epsilon_i,
\epsilon_i \sim N(0, \sigma^2), with balanced arms:
\text{Var}(\hat\beta) = \frac{4\sigma^2}{n_H}
Value
Numeric. Effective information.
Examples
d_eff_continuous(n_sg = 100, sigma_y = 1.5)
# 44.4
Effective Information for Count Outcomes (Poisson)
Description
Computes d_{\text{eff}} = D / \phi where D is the total
expected events and \phi is the dispersion parameter.
Usage
d_eff_count(
n_sg = NULL,
event_rate = NULL,
total_events = NULL,
dispersion = 1
)
Arguments
n_sg |
Integer. Subgroup sample size. Required if
|
event_rate |
Numeric. Mean events per patient. Required if
|
total_events |
Numeric. Total observed events |
dispersion |
Numeric. Overdispersion parameter (1.0 for Poisson, greater than 1 for quasi-Poisson). Default: 1.0. |
Details
For Poisson regression with log-offset:
\text{Var}(\hat\beta) = \phi \left(\frac{1}{D_0} + \frac{1}{D_1}\right)
\approx \frac{4\phi}{D}
Under proportional hazards, D = d (Cox events), so
d_{\text{eff}}^{\text{Poisson}} = d_{\text{eff}}^{\text{Cox}}
(Laird and Olivier, 1981).
Value
Numeric. Effective information.
Examples
# From total events
d_eff_count(total_events = 55)
# From sample size and rate
d_eff_count(n_sg = 100, event_rate = 0.8)
# Quasi-Poisson with overdispersion
d_eff_count(total_events = 55, dispersion = 1.5)
Effective Information for Survival (Cox PH)
Description
Computes d_{\text{eff}} = n_H (1 - p_c) where p_c is the
censoring proportion.
Usage
d_eff_survival(n_sg, prop_cens)
Arguments
n_sg |
Integer. Subgroup sample size. |
prop_cens |
Numeric. Proportion censored (0–1). |
Value
Numeric. Effective information (expected events).
Examples
d_eff_survival(n_sg = 60, prop_cens = 0.45)
# 33
Default GRF parameters (general)
Description
Default GRF parameters (general)
Usage
default_grf_params_gen()
Default ForestSearch parameters (general)
Description
Returns the default fs_params list used by
run_simulation_analysis(). The parallel_args entry
defaults to list(plan = "sequential") because
run_simulation_analysis() is designed to be called inside a
foreach() %dofuture% loop (one replicate per worker). Running
the inner forestsearch() multisession in that context produces
nested parallelism: each outer worker tries to spawn its own pool of
inner workers, which parallelly rejects with a 300\
limit error. Users who call run_simulation_analysis() once at
the top level (i.e., not inside %dofuture%) can opt back into
multisession by passing
fs_params = list(parallel_args = list(plan = "multisession", workers = N)).
Usage
default_sim_params()
Define Subgroups with Flexible Cutpoints
Description
Define Subgroups with Flexible Cutpoints
Usage
define_subgroups(
df_work,
data,
subgroup_vars,
subgroup_cuts,
continuous_vars,
model,
verbose
)
Bivariate Density for Split-Sample HR Threshold Detection
Description
Computes the joint density for the two-split detection criterion where both split-halves must exceed individual thresholds and their average must exceed a consistency threshold.
Usage
density_threshold_both(x, theta, prop_cens = 0.3, n_sg, k_avg, k_ind)
Arguments
x |
Numeric vector of length 2. Log hazard ratio estimates from the two split-halves. |
theta |
Numeric. True hazard ratio in the subgroup. |
prop_cens |
Numeric. Proportion censored (0-1). Default: 0.3 |
n_sg |
Integer. Subgroup sample size. |
k_avg |
Numeric. Threshold for average log(HR) across splits. Typically log(hr.threshold). |
k_ind |
Numeric. Threshold for individual split log(HR). Typically log(hr.consistency). |
Details
The detection criterion requires:
Average of two splits: (x1 + x2)/2 >= k_avg
Individual splits: x1 >= k_ind AND x2 >= k_ind
Under the Leon et al. (2024, Section 2.1) approximation, each 50/50 split-half log(HR) estimator follows N(log(theta), 8/d) where d = n_sg * (1 - prop_cens) is the expected TOTAL number of events in the subgroup. Each split holds d/2 events, so its variance is 4/(d/2) = 8/d.
Value
Numeric. Joint density value at x, or 0 if thresholds not met.
Vectorized Density for Integration
Description
Wrapper around density_threshold_both for use with cubature integration.
Usage
density_threshold_integrand(x, theta, prop_cens, n_sg, k_avg, k_ind)
Arguments
x |
Numeric vector of length 2. Log hazard ratio estimates from the two split-halves. |
theta |
Numeric. True hazard ratio in the subgroup. |
prop_cens |
Numeric. Proportion censored (0-1). Default: 0.3 |
n_sg |
Integer. Subgroup sample size. |
k_avg |
Numeric. Threshold for average log(HR) across splits. Typically log(hr.threshold). |
k_ind |
Numeric. Threshold for individual split log(HR). Typically log(hr.consistency). |
Value
Numeric density value.
Automatically Detect Variable Types in a Dataset
Description
Analyzes a data frame to automatically classify variables as continuous or categorical, and returns a subset of the data with specified variables excluded.
Usage
detect_variable_types(data, max_unique_for_cat = 10, exclude_vars = NULL)
Arguments
data |
A data frame to analyze |
max_unique_for_cat |
Integer. Maximum number of unique values for a numeric variable to be considered categorical. Default is 10. |
exclude_vars |
Character vector of variable names to exclude from both classification and the returned dataset (e.g., ID variables, timestamps). Default is NULL. |
Details
The function classifies variables using the following rules:
Numeric variables with more than
max_unique_for_catunique values are classified as continuousNumeric variables with
max_unique_for_cator fewer unique values are classified as categoricalFactor, character, and logical variables are always classified as categorical
Variables listed in
exclude_varsare omitted from classification and removed from the returned dataset
Value
A list containing:
continuous_vars |
Character vector of variable names classified as continuous |
cat_vars |
Character vector of variable names classified as categorical |
data_subset |
Data frame with exclude_vars columns removed |
Examples
## Not run:
example_data <- data.frame(
id = 1:100,
age = rnorm(100, 50, 10),
grade = sample(1:3, 100, replace = TRUE),
status = sample(c("Active", "Inactive"), 100, replace = TRUE),
score = runif(100, 0, 100)
)
result <- detect_variable_types(example_data,
max_unique_for_cat = 10,
exclude_vars = "id")
result$continuous_vars # c("age", "score")
result$cat_vars # c("grade", "status")
names(result$data_subset) # c("age", "grade", "status", "score")
## End(Not run)
Estimate heterogeneous treatment effects via DINA (data-frame interface)
Description
dina() is the high-level user-facing wrapper around dina_fit().
It accepts a data frame plus column-name arguments, matching the
conventions used elsewhere in the forestsearch package, and
extracts the covariate matrix, response, and treatment indicator
before dispatching to dina_fit().
Usage
dina(
df,
outcome,
treatment,
covariates,
family = c("gaussian", "binomial", "poisson", "cox"),
status = NULL,
...
)
Arguments
df |
data frame (or tibble / |
outcome |
character(1); the name of the response column in
|
treatment |
character(1); the name of the binary treatment
indicator column in |
covariates |
character vector of length |
family |
one of |
status |
character(1) or |
... |
further arguments forwarded to |
Details
All other arguments are forwarded to dina_fit() via ...,
including propensity_method, baseline_method, cross_fitting,
n_folds, cens_type, cens_params, n_grid, eps, and seed.
See dina_fit() for their full documentation.
Covariate columns must be numeric (or integer). Convert factors to
dummy variables (e.g., with stats::model.matrix()) before calling
if needed. Rows with any NA value in the referenced columns are
rejected with an error; use stats::na.omit() or impute beforehand.
Value
An object of class "dina", as returned by dina_fit().
The $call component is overwritten with the original dina()
call so that print() and summary() show the high-level
invocation rather than the internal matrix-based dispatch.
See Also
dina_fit() for the underlying matrix-based estimator;
dina_bagged() for the bagged version with infinitesimal-jackknife
variance.
Examples
set.seed(1)
n <- 400
df_demo <- data.frame(
y = NA,
w = stats::rbinom(n, 1, 0.5),
x1 = stats::runif(n, -1, 1),
x2 = stats::runif(n, -1, 1),
x3 = stats::runif(n, -1, 1)
)
tau_x <- 0.5 + 0.8 * df_demo$x1 - 0.3 * df_demo$x2
df_demo$y <- 1 + df_demo$x1 + df_demo$w * tau_x + stats::rnorm(n)
fit <- dina(
df = df_demo,
outcome = "y",
treatment = "w",
covariates = c("x1", "x2", "x3"),
family = "gaussian",
seed = 1L
)
coef(fit)
confint(fit)
Bagged DINA with IJ variance (data-frame interface)
Description
dina_bagged() is the high-level user-facing wrapper around
dina_fit_bagged(). It accepts a data frame plus column-name
arguments and dispatches to dina_fit_bagged(), which fits the
DINA estimator on n_bags bootstrap replicates and returns the
infinitesimal-jackknife variance per Wager, Hastie and Efron (2014).
Usage
dina_bagged(
df,
outcome,
treatment,
covariates,
family = c("gaussian", "binomial", "poisson", "cox"),
status = NULL,
...
)
Arguments
df |
data frame (or tibble / |
outcome |
character(1); the name of the response column in
|
treatment |
character(1); the name of the binary treatment
indicator column in |
covariates |
character vector of length |
family |
one of |
status |
character(1) or |
... |
further arguments forwarded to |
Details
All bagging-related arguments (n_bags, cross_fitting_per_bag,
n_folds_per_bag, parallel, ij_finite_sample_correction,
project_psd, verbose, ...) are forwarded via .... See
dina_fit_bagged() for full documentation.
For parallel = "bags", set future::plan() to a parallel backend
before calling dina_bagged(). Restore with
future::plan("sequential"); gc() afterwards.
Value
An object of class c("dina_bagged", "dina"), as returned
by dina_fit_bagged(). The $call component is overwritten with
the original dina_bagged() call.
See Also
dina_fit_bagged() for the underlying matrix-based bagged
estimator; dina() for the single-pass DINA with sandwich variance.
Examples
set.seed(1)
n <- 400
df_demo <- data.frame(
y = NA,
w = stats::rbinom(n, 1, 0.5),
x1 = stats::runif(n, -1, 1),
x2 = stats::runif(n, -1, 1),
x3 = stats::runif(n, -1, 1)
)
tau_x <- 0.5 + 0.8 * df_demo$x1 - 0.3 * df_demo$x2
df_demo$y <- 1 + df_demo$x1 + df_demo$w * tau_x + stats::rnorm(n)
fit <- dina_bagged(
df = df_demo,
outcome = "y",
treatment = "w",
covariates = c("x1", "x2", "x3"),
family = "gaussian",
n_bags = 30L,
seed = 1L
)
coef(fit)
sqrt(diag(vcov(fit)))
Estimate heterogeneous treatment effects via DINA
Description
dina_fit() implements the two-step DINA (DIfference in NAtural parameters)
estimator of Gao and Hastie (2025) for heterogeneous treatment effects
(HTE) with general response types: Gaussian (mean difference), binomial
(log odds ratio), Poisson (log mean ratio), and Cox (log hazard ratio).
The treatment effect is assumed to be linear in the covariates,
\tau(x) = \beta_0 + x^\top \beta_{1:d},
following Proposition 1 of the paper. The procedure has two steps:
(1) estimate nuisance functions a(x) and \nu(x) on a training
fold, and (2) plug those estimators into a GLM (or Cox) likelihood on the
held-out fold to estimate \beta. Cross-fitting averages the two
estimators obtained by swapping the folds.
Usage
dina_fit(
X,
Y,
W,
family = c("gaussian", "binomial", "poisson", "cox"),
propensity_method = "logistic",
baseline_method = NULL,
cross_fitting = TRUE,
n_folds = 2L,
cens_type = NULL,
cens_params = list(),
n_grid = 500L,
eps = 1e-06,
seed = NULL
)
Arguments
X |
a numeric covariate matrix or data frame with |
Y |
the response. For |
W |
binary treatment indicator (0 = control, 1 = treatment), length
|
family |
character; one of |
propensity_method |
how to estimate the propensity score
|
baseline_method |
how to estimate the natural-parameter functions
|
cross_fitting |
logical; if |
n_folds |
positive integer giving the number of cross-fitting folds
(default |
cens_type |
for |
cens_params |
named list of censoring parameters. Required entries
depend on
|
n_grid |
positive integer giving the number of time-grid points used
for numerical integration of the not-censoring probability for
|
eps |
small positive constant added to denominators for numerical
stability (default |
seed |
optional integer seed for the random cross-fitting fold assignment. When supplied, the global RNG state is restored on exit. |
Details
The Gaussian family uses the simplification of Section 3.2: a(x) =
e(x) and \nu(x) = E[Y \mid X = x], so the conditional outcome means
\eta_0(x), \eta_1(x) are not required separately. For the
binomial and Poisson families, separate within-arm fits to \eta_0(x)
and \eta_1(x) are used to construct a(x) via the modified
propensity in Eq. (12) of the paper. For the Cox family, a(x) uses
the not-censoring probabilities of Eq. (18); how those are estimated is
controlled by cens_type.
Variance estimation follows the standard sandwich form:
\widehat{V}(\hat\beta) = \widehat\Sigma_1^{-1} \widehat\Sigma_2
\widehat\Sigma_1^{-1} / n, where \widehat\Sigma_1 and
\widehat\Sigma_2 are the bread (negative score derivative) and meat
(outer product of scores) matrices. Under cross-fitting these are
averaged across folds before the sandwich is formed. For the Cox step 2
we use the partial-likelihood information-matrix variance returned by
survival::coxph().
Value
An object of class "dina", a list with components:
- coefficients
named numeric vector of estimated DINA coefficients
\hat\beta, with(Intercept)first followed by one entry per covariate.- vcov
estimated variance-covariance matrix of
\hat\beta.- family
the response family, as supplied.
- n, d
sample size and number of covariates.
- cross_fitting, n_folds
cross-fitting settings.
- call
the matched call.
References
Gao, Z. and Hastie, T. (2025). Estimating heterogeneous treatment effects for general responses. Biometrics 81(4): ujaf162. doi:10.1093/biomtc/ujaf162.
Robinson, P. M. (1988). Root-n-consistent semiparametric regression. Econometrica 56(4): 931–954.
Nie, X. and Wager, S. (2021). Quasi-oracle estimation of heterogeneous treatment effects. Biometrika 108(2): 299–319.
Dandl, S., Bender, A. and Hothorn, T. (2024). Heterogeneous treatment effect estimation for observational data using model-based forests. Statistical Methods in Medical Research 33(3): 392–413.
Examples
set.seed(1)
n <- 400
d <- 3
X <- matrix(stats::runif(n * d, -1, 1), n, d)
colnames(X) <- paste0("x", seq_len(d))
W <- stats::rbinom(n, 1, stats::plogis(0.5 * X[, 1]))
tau_x <- 0.5 + 0.8 * X[, 1] - 0.3 * X[, 2]
Y <- 1 + X[, 1] + W * tau_x + stats::rnorm(n)
fit <- dina_fit(X, Y, W, family = "gaussian", seed = 1)
coef(fit)
confint(fit)
Bagged DINA with infinitesimal-jackknife variance
Description
dina_fit_bagged() fits the DINA estimator on n_bags bootstrap
replicates of the data and aggregates the resulting per-bag coefficient
vectors into a bagged estimator. The variance is estimated by the
infinitesimal jackknife (IJ) of Wager, Hastie and Efron (2014), with
the finite-sample-corrected form
\widehat{V}_{IJ-U} = \widehat{V}_{IJ} - (n / B) \widehat{V}_{boot},
where \widehat{V}_{boot} is the empirical covariance of the
per-bag coefficients. Negative eigenvalues introduced by the
Monte-Carlo correction (an artifact of small n_bags) are optionally
clipped to zero via PSD projection so that downstream confint() is
always well-defined.
Usage
dina_fit_bagged(
X,
Y,
W,
family = c("gaussian", "binomial", "poisson", "cox"),
propensity_method = "logistic",
baseline_method = NULL,
n_bags = 100L,
cross_fitting_per_bag = TRUE,
n_folds_per_bag = 2L,
cens_type = NULL,
cens_params = list(),
n_grid = 500L,
ij_finite_sample_correction = TRUE,
project_psd = TRUE,
eps = 1e-06,
seed = NULL,
parallel = c("none", "bags"),
verbose = FALSE
)
Arguments
X |
a numeric covariate matrix or data frame with |
Y |
the response. For |
W |
binary treatment indicator (0 = control, 1 = treatment), length
|
family |
character; one of |
propensity_method |
how to estimate the propensity score
|
baseline_method |
how to estimate the natural-parameter functions
|
n_bags |
positive integer giving the number of bootstrap bags
(default |
cross_fitting_per_bag |
logical; if |
n_folds_per_bag |
positive integer giving the number of
cross-fitting folds within each bag (default |
cens_type |
for |
cens_params |
named list of censoring parameters. Required entries
depend on
|
n_grid |
positive integer giving the number of time-grid points used
for numerical integration of the not-censoring probability for
|
ij_finite_sample_correction |
logical; if |
project_psd |
logical; if |
eps |
small positive constant added to denominators for numerical
stability (default |
seed |
optional integer seed for the random cross-fitting fold assignment. When supplied, the global RNG state is restored on exit. |
parallel |
character, one of |
verbose |
logical; if |
Details
The implementation calls dina_fit() once per bag, with cross-fitting
optionally enabled within each bag. All other arguments
(propensity_method, baseline_method, cens_type, cens_params,
n_grid, eps) are passed through unchanged.
Value
An object of class c("dina_bagged", "dina"), a list
containing all the components of a "dina" fit (so that the standard
S3 methods work via inheritance) plus:
- vcov
the corrected, PSD-projected IJ variance matrix (or the raw IJ if
project_psd = FALSE).- vcov_ij_raw
the IJ variance without the finite-sample correction.
- vcov_ij_corrected
the IJ variance with the finite-sample correction, before PSD projection.
- vcov_boot
the empirical bootstrap covariance of the per-bag coefficients,
\widehat{V}_{boot}.- bag_coefficients
n_bags_usedby(d + 1)matrix of per-bag coefficient vectors.- bag_inclusion
nbyn_bags_usedinteger matrix of bag inclusion counts,N_{bi}^*.- n_bags, n_bags_used
requested and successfully-fit bag counts.
- n_folds_per_bag, cross_fitting_per_bag
bag-level cross-fitting settings.
References
Wager, S., Hastie, T. and Efron, B. (2014). Confidence intervals for random forests: the jackknife and the infinitesimal jackknife. Journal of Machine Learning Research 15: 1625–1651.
Gao, Z. and Hastie, T. (2025). Estimating heterogeneous treatment effects for general responses. Biometrics 81(4): ujaf162. doi:10.1093/biomtc/ujaf162.
See Also
dina_fit() for the single-pass DINA estimator with
sandwich variance.
Examples
set.seed(1)
n <- 400
d <- 3
X <- matrix(stats::runif(n * d, -1, 1), n, d)
colnames(X) <- paste0("x", seq_len(d))
W <- stats::rbinom(n, 1, stats::plogis(0.5 * X[, 1]))
tau_x <- 0.5 + 0.8 * X[, 1] - 0.3 * X[, 2]
Y <- 1 + X[, 1] + W * tau_x + stats::rnorm(n)
fit <- dina_fit_bagged(X, Y, W, family = "gaussian",
n_bags = 30L, seed = 1L)
coef(fit)
sqrt(diag(vcov(fit)))
Extract DINA per-covariate Pareto frontiers as forestsearch cuts
Description
Runs the same univariate (covariate, direction, threshold) search as
dina_subgroup(), but instead of selecting one subgroup it returns a
pool of per-covariate Pareto-non-dominated cuts. Each cut is
rendered as a forestsearch cut expression ("x1 <= 0.5"), so the pool
can be fed to forestsearch() as screening-stage candidates in the
same way GRF tree cuts are. This is the DINA analog of GRF candidate
discovery.
Usage
dina_frontier(
fit,
df,
covariates,
scope = c("wide", "harm"),
m_diff = NULL,
n_min = 60L,
direction = c("both", "left", "right"),
max_per_covariate = Inf,
max_subgroups = Inf,
digits = 3L
)
Arguments
fit |
a fitted DINA object (class |
df |
data frame containing the covariate columns. |
covariates |
character vector of numeric covariate columns to search over. |
scope |
one of |
m_diff |
scalar harm floor on the link (natural-parameter) scale.
Required when |
n_min |
positive integer minimum subgroup size. Default |
direction |
one of |
max_per_covariate |
positive integer, or |
max_subgroups |
positive integer, or Both caps default to |
digits |
significant figures for rounding the emitted threshold
in |
Details
Why per-covariate and not a single global frontier. The global Pareto frontier in (effect, N) space collapses onto whichever single covariate dominates the trade-off boundary, returning many micro-stepped cuts on that one covariate. That is faithful to the definition but the opposite of what forestsearch needs: its value is in composing cuts drawn from several covariates. So this function computes a separate frontier for each covariate, then pools them, mirroring how a GRF tree splits on multiple variables.
Two caps, each optional. Both caps act on this function's return
value and nothing else: they trim the frontier table of single cuts
that is returned here. They are not a cap on a selection family – the
rows they trim are univariate cuts, never conjunctions – and
dina_subgroup() does not take them, so nothing they do reaches the
subgroup selector. Either may be Inf to disable it:
-
max_per_covariatebounds how many cuts any single covariate contributes (its top cuts by effect), limiting within-covariate redundancy. -
max_subgroupsbounds the total pool size across covariates.
Under forestsearch(subgroup_method = "dina") neither cap has any effect:
that path selects with dina_subgroup() and never consults a frontier
table. They bite only where this function's output is used directly –
including forestsearch(use_dina = TRUE, dina_args = list(selected_only = FALSE)), where cut_expr becomes the screening-stage candidate pool.
When the pool exceeds max_subgroups, the global trim is round-robin
on within-covariate effect rank – every covariate's best cut first
(ordered by effect), then every covariate's second-best, and so on –
so a finite budget preserves cross-covariate spread instead of
refilling with one covariate's micro-steps. Setting
max_per_covariate = Inf recovers a pure global budget; setting
max_subgroups = Inf recovers per-covariate-only limits. Both default to
Inf, and a finite cap that actually removes rows warns rather than
trimming the display silently.
Cut form. The expression is always the canonical "<covariate> <= <threshold>". forestsearch's factor machinery exposes both the
cut and its complement, so one expression covers both of DINA's
left/right directions – the direction column is retained for
reference only. Thresholds are rounded to digits significant
figures so they deduplicate cleanly against existing GRF / median cuts
and read well; raw values remain in threshold.
These cuts are candidates, not forced selections: forestsearch
composes (up to maxk) and consistency-gatekeeps them, exactly as it
does GRF cuts. To use them without disturbing any user-supplied
forced cuts, append rather than overwrite, e.g.
forestsearch(..., conf_force = c(my_forced_cuts, fr$cut_expr)).
Value
A data frame (one row per cut, after the caps and threshold-rounding deduplication) with columns:
- covariate, direction, threshold
the subgroup definition (raw threshold).
- n_subgroup
subgroup size.
- mean_tau_hat
subgroup-mean tau-hat on the link scale.
- effect
the same on the natural scale (exponentiated for ratio families binomial/poisson/cox), used for the frontiers and the effect ranking.
- cut_expr
the canonical forestsearch cut expression.
The result carries attributes n_frontier (total per-covariate
frontier size, deduped, before the caps), n_qualifying, and
scope. Returns a 0-row data frame (same columns) when no candidate
qualifies.
See Also
dina_subgroup() for the single-subgroup selector;
forestsearch() for the screening stage that consumes the cuts;
compute_pareto_frontier() for the post-hoc frontier reporting on a
forestsearch result.
Examples
set.seed(1)
n <- 400
df_demo <- data.frame(
w = stats::rbinom(n, 1, 0.5),
x1 = stats::runif(n, -1, 1),
x2 = stats::runif(n, -1, 1),
x3 = stats::runif(n, -1, 1)
)
tau_x <- 0.3 + 1.2 * df_demo$x1 - 0.4 * df_demo$x2
df_demo$y <- 0.5 * df_demo$x1 + df_demo$w * tau_x + stats::rnorm(n)
fit <- dina(df_demo, outcome = "y", treatment = "w",
covariates = c("x1", "x2", "x3"),
family = "gaussian", seed = 1L)
fr <- dina_frontier(fit, df_demo, covariates = c("x1", "x2", "x3"),
n_min = 60L)
fr
fr$cut_expr # ready to append to forestsearch(conf_force = ...)
Identify a harm subgroup from a DINA fit via univariate threshold search
Description
Given a fitted DINA object ("dina" or "dina_bagged") and a data
frame containing the covariates the model was fit on, search over
(covariate, threshold, direction) candidates for subgroups whose mean
per-patient HTE estimate meets the harm threshold m_diff, then select
one according to sg_focus (default "maxSG": the largest qualifying
subgroup). This is the simplest bridge between DINA's continuous
estimator \hat\tau(x) = \hat\beta_0 + x'\hat\beta and
forestsearch's interpretable signature framework.
Usage
dina_subgroup(
fit,
df,
covariates,
m_diff,
n_min = 60L,
n_min.frac = 0.1,
direction = c("both", "left", "right"),
max_depth = 2L,
grid_probs = seq(0.1, 0.9, by = 0.1),
sg_focus = "maxSG",
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
alpha = 0.05,
tau_sign = 1
)
Arguments
fit |
a fitted DINA object of class |
df |
data frame containing the covariate columns referenced by
|
covariates |
character vector of column names in |
m_diff |
scalar harm threshold on the natural-parameter scale
of It is a candidacy floor, not a ranking key. It decides which
candidates enter the qualifying family; Scale. The link (tau) scale of Under Under |
n_min |
positive integer, or |
n_min.frac |
numeric in (0, 1); fraction of |
direction |
one of |
max_depth |
integer, |
grid_probs |
numeric vector of probabilities in |
sg_focus |
character; subgroup selection criterion applied to the
qualifying candidates. One of Five spellings, one rule. On this path
|
selection_rule |
character; rule defining the candidate inclusion
band for |
effect_neighborhood |
numeric in |
alpha |
confidence level for the Wald interval on the
subgroup-mean tau-hat. Default |
tau_sign |
|
Details
For each covariate x_j in covariates, the search visits the
sorted unique values of df[[x_j]] as candidate thresholds q and
for each threshold considers the left-tail subgroup of patients with
x_{j,i} \le q (and optionally also the right-tail subgroup
with x_{j,i} \ge q). A candidate
subgroup must satisfy:
subgroup size
|S| >= n_min, ANDmean per-patient tau-hat over
S>= m_diff.
Among the qualifying candidates, sg_focus selects the returned
subgroup, mirroring forestsearch()'s selection vocabulary:
-
"maxSG"(default) – largest qualifying subgroup (legacy behaviour; size ties broken by the most extreme mean tau-hat). -
"minSG"– smallest qualifying subgroup. -
"eff"– most extreme effect, ignoring size. -
"effMaxSG"– largest subgroup in the candidate inclusion band (seeselection_rule). -
"effMinSG"– smallest subgroup in the inclusion band. The"eff*"band foci reuse the inclusion-band logic shared withsort_subgroups(); for ratio families (binomial, poisson, cox) the band is applied on the natural (exponentiated) effect scale.selection_ruledefines the band, exactly as inforestsearch():"neighborhood"(1-D effect band withineffect_neighborhoodof the maximum qualifying effect),"pareto"(the 2-D non-dominated frontier in (effect, N) space), or"both"(their intersection). DINA's univariate (covariate, threshold, direction) candidates populate the (effect, N) plane just as forestsearch's signatures do, so the frontier is defined analogously.
Value
An object of class "dina_subgroup", a list with components:
- found
logical;
TRUEif any candidate satisfiedm_diffandn_min, otherwiseFALSE. WhenFALSE, the subgroup-specific components below areNULL.- covariate, direction, threshold
the chosen cut(s). Length 1 for a single-covariate subgroup; length 2 (one entry per covariate) when
max_depth = 2Lselects a conjunction.- labels
character vector of the chosen
{var op q}factor label(s), in the form consumed byget_dfpred()(AND-composed). Length matchescovariate.- label
a single combined display string, e.g.
"{nodes >= 10 & age >= 60}".- depth
integer; number of covariates in the selected subgroup (
1or2).- n_subgroup
integer size of the chosen subgroup.
- mean_tau_hat
scalar subgroup-mean tau-hat (in the
tau_signorientation).- se_mean_tau_hat
Wald standard error, computed as
sqrt(a_S^T vcov(fit) a_S)(CONDITIONAL on the chosen subgroup – not selection-adjusted).- ci
named length-2 vector
(lower, upper)giving the Wald 1 - alpha CI.- mask
logical vector of length
nrow(df)marking which rows are in the chosen subgroup.- m_diff, n_min, sg_focus, selection_rule, effect_neighborhood, alpha, family, max_depth
the inputs / fit family, echoed for reproducibility.
- n_total
nrow(df).- n_candidates_searched
total number of
(j, dir, q)triples evaluated.- n_candidates_qualifying
number of those triples that met both the size and
m_diffconstraints.- call
the matched call.
Depth-2 conjunctions
With max_depth = 2 the candidate set is extended to AND-conjunctions
of two thresholds on distinct covariates, e.g.
nodes >= q1 & age >= q2. Because tau-hat is linear in X, the mean
over the intersection mask is scored exactly like a depth-1 cut, so
the m_diff, n_min, and sg_focus machinery is unchanged. To stay
tractable, depth-2 pairs are generated over a per-covariate quantile
grid (grid_probs) rather than every unique value, while depth-1
singletons retain full resolution. Singletons and pairs are ranked
jointly under sg_focus, so a depth-2 conjunction is returned only
when it outranks every depth-1 candidate; otherwise the depth-1
selection is recovered exactly. When the selected subgroup is a
conjunction, covariate, direction, and threshold are
length-2 vectors and labels holds the two {var op q} factors in
the form consumed by get_dfpred().
The variance of the subgroup-mean tau-hat is computed as
\widehat{\mathrm{Var}}(\bar{\hat\tau}_S) =
a_S^\top \widehat{\mathrm{Var}}(\hat\beta) a_S,
with a_S = (1, \bar{x}_S^\top) and
\widehat{\mathrm{Var}}(\hat\beta) = \mathrm{vcov}(\mathtt{fit}).
The variance source therefore tracks the fit class: sandwich for
"dina" and infinitesimal jackknife for "dina_bagged".
Inference caveat
The returned Wald confidence interval is conditional on the
selected (covariate, threshold) being treated as pre-specified.
It does not adjust for the selection across the candidate set and
will undercover when the search space is large. Use this CI for
descriptive reporting; for hypothesis testing against m_diff,
a selection-adjusted interval via bootstrap of the full search
procedure is required. The undercoverage is more pronounced with
max_depth = 2L, whose candidate space is substantially larger.
See Also
dina_fit() / dina() for the underlying estimator;
dina_fit_bagged() / dina_bagged() for the bagged variant whose
IJ variance propagates through the subgroup-mean SE here;
sort_subgroups() and forestsearch() for the selection
vocabulary mirrored by sg_focus.
Examples
set.seed(1)
n <- 400
df_demo <- data.frame(
w = stats::rbinom(n, 1, 0.5),
x1 = stats::runif(n, -1, 1),
x2 = stats::runif(n, -1, 1),
x3 = stats::runif(n, -1, 1)
)
tau_x <- 0.3 + 1.2 * df_demo$x1 - 0.4 * df_demo$x2
df_demo$y <- 0.5 * df_demo$x1 + df_demo$w * tau_x + stats::rnorm(n)
fit <- dina(df_demo, outcome = "y", treatment = "w",
covariates = c("x1", "x2", "x3"),
family = "gaussian", seed = 1L)
sg <- dina_subgroup(fit, df_demo,
covariates = c("x1", "x2", "x3"),
m_diff = 0.5, n_min = 60L)
sg
# Tighter subgroup via the effect-neighborhood band:
dina_subgroup(fit, df_demo, covariates = c("x1", "x2", "x3"),
m_diff = 0.5, n_min = 60L, sg_focus = "effMaxSG")
# Frontier-aware selection (2-D non-dominated set in (effect, N)):
dina_subgroup(fit, df_demo, covariates = c("x1", "x2", "x3"),
m_diff = 0.5, n_min = 60L, sg_focus = "effMaxSG",
selection_rule = "both")
# Depth-2: allow AND-conjunctions of two covariates.
dina_subgroup(fit, df_demo, covariates = c("x1", "x2", "x3"),
m_diff = 0.5, n_min = 60L, max_depth = 2L)
Selection-adjusted bootstrap CI for dina_subgroup()
Description
Performs selection-adjusted inference for the harm subgroup discovered
by dina_subgroup(). For each of n_boot bootstrap resamples of the
input data frame, the function refits DINA via dina() and re-runs
the subgroup search via dina_subgroup(), then reports three classes
of output:
Usage
dina_subgroup_bootstrap(
df,
outcome,
treatment,
covariates,
family = c("gaussian", "binomial", "poisson", "cox"),
status = NULL,
m_diff,
n_min = 60L,
n_min.frac = 0.1,
direction = c("both", "left", "right"),
sg_focus = "maxSG",
effect_neighborhood = 0.1,
alpha = 0.05,
n_boot = 200L,
parallel = c("none", "boots"),
seed = NULL,
verbose = FALSE,
fit = NULL,
sg = NULL,
refit = TRUE,
refit_strata = NULL,
refit_confounders = "none",
...
)
Arguments
df |
data frame containing the outcome, treatment, and covariate columns. |
outcome |
character(1); name of the response column in |
treatment |
character(1); name of the binary treatment indicator column. |
covariates |
character vector; names of the covariate columns to search over. All must be numeric. |
family |
one of |
status |
character(1) or |
m_diff |
scalar harm threshold on the natural-parameter scale. |
n_min |
positive integer, or |
n_min.frac |
numeric in (0, 1); fraction of |
direction |
one of |
sg_focus |
character; subgroup selection criterion, forwarded to
|
effect_neighborhood |
numeric in |
alpha |
confidence level for the percentile CI. Default |
n_boot |
positive integer; number of bootstrap iterations.
Default |
parallel |
one of |
seed |
optional integer seed for reproducibility of the
bootstrap indices. Default |
verbose |
logical; if |
fit |
optional precomputed DINA fit ( |
sg |
optional precomputed |
refit |
logical; if |
refit_strata |
|
refit_confounders |
adjustment set for the within-subgroup
model: |
... |
further arguments forwarded to |
Details
An effect CI on
a*^T beta_b, wherea* = (1, x_bar_S)is held FIXED from the original-data subgroupS*andbeta_bis each iteration's DINA coefficient vector. This is the bootstrap analog of the Wald CI returned bydina_subgroup()– conditional on the discovered subgroup definition, with width reflecting sampling variance ofbeta-hat.-
Stability CIs on the bootstrap-selected
n_subgroupandthreshold, restricted to iterations that selected the same(covariate, direction)as the original-data subgroup. A selection-frequency table counting how often each
(covariate, direction)was chosen. The covariate-selection stability diagnostic.
Each bootstrap iteration calls dina() (single-pass, sandwich
variance) under the hood, not dina_bagged().
If you have already run dina() and dina_subgroup() on the original
data, pass those objects via fit and/or sg so the original-data
point estimate reuses them rather than being recomputed here. This
guarantees the reported point estimate matches your standalone result
exactly; otherwise the internal refit uses its own cross-fitting fold
assignment, which at a boundary-case m_diff can disagree with the
upstream fit on whether a subgroup qualifies.
Value
An object of class "dina_subgroup_bootstrap", a list with
components:
- point
the
"dina_subgroup"object computed on the original (unbootstrapped) data.- effect_ci
named length-2 numeric vector (
lower,upper) giving the bootstrap percentile CI on the fixed-subgroup linear functionala*^T beta_b, wherea* = (1, x_bar_S)is the covariate-mean vector of the original-data subgroup andbeta_bis each bootstrap iteration's DINA coefficient vector. This is the meaningful effect-uncertainty CI; width reflects sampling variance ofbeta-hat.- effect_dist
numeric vector of length
n_bootcontaininga*^T beta_bfor each iteration (NAfor iterations whose DINA fit failed).- refit
the
"dina_subgroup_refit"object (standard within-subgroup model on the original data), orNULLifrefit = FALSE, no subgroup was found, or the fit failed.- refit_effect_ci
named length-2 numeric vector (
lower,upper); bootstrap percentile CI on the within-subgroup standard-model treatment effect, refit within the FIXED discovered signature on each resample. Conditional on the signature; comparable toeffect_ci.(NA, NA)ifrefitwas not performed.- refit_effect_dist
numeric vector of length
n_bootof per-iteration within-signature standard-model treatment effects (NAwhere not computed or the fit failed).- n_subgroup_ci, threshold_ci
named length-2 numeric vectors giving percentile CIs for the structural quantities of bootstrap-selected subgroups, restricted to iterations that selected the same
(covariate, direction)as the original-data subgroup. Subgroup-size and boundary-location stability diagnostics.- n_modal_match
integer; number of bootstrap iterations whose selection matched the original-data
(covariate, direction). Denominator for the stability CIs.- selection_frequency
a
tableof(covariate, direction)selection counts across found iterations. Covariate-selection stability diagnostic.- boot_results
data frame with one row per bootstrap iteration; columns
found,covariate,direction,threshold,n_subgroup,mean_tau_hat,failed. Themean_tau_hatcolumn records the bootstrap-selected subgroup's boundary value (pinned nearm_diffby the search rule) and is retained for diagnostic transparency rather than as a meaningful effect estimate.- n_boot, n_boot_found, n_boot_failed
convergence diagnostics.
- alpha, m_diff, n_min, sg_focus, effect_neighborhood, family
echoed inputs.
- call
the matched call.
Why the bootstrap target is the fixed-subgroup effect
A naive bootstrap of dina_subgroup()'s mean_tau_hat is
uninformative. Under the default sg_focus = "maxSG" the search
selects the largest subgroup whose mean tau-hat exceeds m_diff,
so the chosen subgroup lies at the boundary by construction: adding
one more patient would push the mean below m_diff. Each bootstrap
iteration's selected subgroup is similarly pinned, so the bootstrap
distribution of mean_tau_hat clusters tightly above m_diff
and the resulting CI quantifies threshold-grid granularity
rather than effect uncertainty. (Under the band foci
"effMaxSG" / "effMinSG" the per-iteration subgroup is instead
anchored to the maximum effect rather than the m_diff boundary,
but the same caveat applies: mean_tau_hat is a selection-pinned
quantity, not an effect estimate.) The fixed-subgroup linear
functional a*^T beta_b instead targets the clinically
meaningful estimand "treatment effect for patients whose
covariates lie in the discovered subgroup". Because DINA's
beta-hat is cross-fit (and hence asymptotically unbiased even
for data-driven contrasts), the percentile distribution of
a*^T beta_b gives a non-parametric CI that complements the
sandwich Wald CI of dina_subgroup().
Two complementary effect estimates
When refit = TRUE (default), the function reports two treatment
effects for the discovered subgroup, on the same natural-parameter
scale:
the DINA effect (
effect_ci): the BLP-analoga*^T beta, a projection of the cross-fitted CATE surface averaged over the subgroup; andthe within-subgroup standard-model effect (
refit_effect_ci): the treatment coefficient from a plain Cox model (or GLM) fit on the subgroup, refit within the FIXED discovered signature on each resample. This is the contrast a clinical-trial reader usually expects.
The two coincide only under correct linearity; reporting both is
informative. Crucially, both CIs are conditional on the
discovered signature treated as pre-specified – neither corrects for
the data-driven selection of the signature. A selection-adjusted
(de-biased) interval requires the more involved bias-corrected
bootstrap of Leon et al. (2024, Section 3.2) and is not provided
here; selection stability is instead summarized by
selection_frequency.
Performance
With parallel = "boots", each bootstrap iteration runs in a worker
via foreach::foreach() %dofuture%. Set future::plan() to a
parallel backend (e.g., future::multisession) before calling, and
reset to "sequential" afterwards. Workers run from the installed
package; devtools::install() is required after code edits before
the parallel branch will see them.
See Also
dina_subgroup() for the underlying threshold search;
dina() for the per-iteration DINA fit;
dina_subgroup_refit for the standard within-subgroup
model reported when refit = TRUE.
Examples
## Not run:
set.seed(1)
n <- 400
df_demo <- data.frame(
w = stats::rbinom(n, 1, 0.5),
x1 = stats::runif(n, -1, 1),
x2 = stats::runif(n, -1, 1)
)
tau_x <- 0.4 + 1.2 * df_demo$x1
df_demo$y <- 0.5 * df_demo$x1 + df_demo$w * tau_x + stats::rnorm(n)
bs <- dina_subgroup_bootstrap(
df = df_demo,
outcome = "y",
treatment = "w",
covariates = c("x1", "x2"),
family = "gaussian",
m_diff = 0.5,
n_boot = 50L,
seed = 1L
)
print(bs)
## End(Not run)
Standard within-subgroup treatment-effect model for a DINA subgroup
Description
Given a subgroup discovered by dina_subgroup, fit the
standard treatment-effect model on the patients in that
subgroup: a Cox proportional-hazards model when
sg$family == "cox", or a generalized linear model otherwise.
The reported effect is the treatment coefficient on the
natural-parameter scale (log-HR for Cox, log-OR for binomial,
log-IRR for Poisson, mean difference for Gaussian), directly
comparable to the DINA effect returned by dina_subgroup() but
obtained from a conventional within-subgroup regression rather than
from the cross-fitted CATE surface.
Usage
dina_subgroup_refit(
sg,
df,
treatment,
outcome,
covariates,
status = NULL,
strata = NULL,
confounders = "none",
alpha = 0.05
)
Arguments
sg |
a |
df |
data frame the subgroup was identified on; must contain the
outcome, treatment, covariate, and any |
treatment |
character(1); name of the binary treatment column. |
outcome |
character(1); name of the response column. For
|
covariates |
character vector of the DINA covariate names. Used
only to construct the automatic adjustment set when
|
status |
character(1) or |
strata |
|
confounders |
|
alpha |
confidence level for the Wald interval. Default
|
Details
The model is unadjusted by default (~ treatment); optional
covariate adjustment and stratification are available via
confounders and strata.
Value
An object of class "dina_subgroup_refit", a list with
components:
- effect
treatment coefficient on the natural-parameter scale (the point estimate).
- se
Wald standard error of
effect.- ci
named length-2 numeric vector (
lower,upper); Wald1 - alphaCI on the natural-parameter scale.- effect_scale
character label for the scale: one of
"log-HR","log-OR","log-IRR","MD".- ratio_scale
logical;
TRUEwheneffectis a log-ratio (Cox/binomial/Poisson) so thatexp()maps it to HR/OR/IRR.- signature
character; the subgroup rule, e.g.
"nodes >= 12".- formula
the fitted model formula.
- confounders_used, strata_used
the resolved adjustment set and stratification columns actually used.
- n_subgroup, n_treated, n_control
subgroup size and per-arm counts.
- n_events, n_events_treated, n_events_control
event counts within the subgroup (Cox only;
NAotherwise).- family, alpha
echoed inputs.
- model
the fitted
coxph/glmobject.- call
the matched call.
Adjustment set (the confounders argument)
Three modes:
-
"none"(default) – unadjusted:~ treatment. -
NULL– automatic: the DINA covariates supplied incovariates, with the subgroup-defining covariate (sg$covariate) omitted. Omitting it avoids adjusting on a variable whose range is restricted by the subgroup definition. a character vector – exactly those columns. A warning is issued if it includes
sg$covariate.
Inference caveat
The Wald confidence interval treats the subgroup definition as pre-specified. It does not adjust for the data-driven selection of the signature; it is a within-signature (conditional) interval, not a selection-adjusted one. A selection-adjusted interval requires a bias-corrected bootstrap of the full search procedure (Leon et al. 2024, Section 3.2) and is not provided here.
See Also
dina_subgroup for the subgroup search and its
DINA (BLP-analog) effect; dina_subgroup_bootstrap for
bootstrap inference that reports this within-subgroup model
alongside the DINA effect.
Examples
## Not run:
fit <- dina(df, outcome = "time", treatment = "trt",
covariates = c("age", "nodes", "er"),
family = "cox", status = "status", seed = 1L)
sg <- dina_subgroup(fit, df, covariates = c("age", "nodes", "er"),
m_diff = 0, n_min = 60L)
# Unadjusted within-subgroup Cox model (default):
dina_subgroup_refit(sg, df, treatment = "trt", outcome = "time",
covariates = c("age", "nodes", "er"),
status = "status")
# Adjusted for DINA covariates except the subgroup-defining one:
dina_subgroup_refit(sg, df, treatment = "trt", outcome = "time",
covariates = c("age", "nodes", "er"),
status = "status", confounders = NULL)
## End(Not run)
Dummy-code a data frame (numeric pass-through, factors expanded)
Description
Dummy-code a data frame (numeric pass-through, factors expanded)
Usage
dummy_encode(df)
Arguments
df |
Data frame with numeric and/or factor columns. |
Value
Data frame with numeric columns unchanged and factor columns
expanded via acm.disjctif.
Examples
df <- data.frame(age = c(40, 55, 70), grade = factor(c("1", "2", "3")))
dummy_encode(df)
Early Stopping Decision
Description
Evaluates whether enough evidence exists to stop early based on confidence interval for consistency proportion.
Usage
early_stop_decision(
n_success,
n_total,
threshold,
conf.level = 0.95,
min_samples = 20
)
Arguments
n_success |
Integer. Number of splits meeting consistency. |
n_total |
Integer. Total number of valid splits. |
threshold |
Numeric. Target consistency threshold. |
conf.level |
Numeric. Confidence level for decision (default 0.95). |
min_samples |
Integer. Minimum samples before allowing early stop. |
Value
Character. One of "continue", "pass", or "fail".
Examples
early_stop_decision(95, 100, threshold = 0.90)
early_stop_decision(60, 100, threshold = 0.90)
early_stop_decision(10, 15, threshold = 0.90) # below min_samples
Human-Readable Effect Measure Label
Description
Returns a descriptive label for use in table footnotes and narrative
text. Falls back to "HR" for survival analyses.
Usage
effect_measure_label(effect_measure = NULL, outcome_type = NULL)
Arguments
effect_measure |
Character or |
outcome_type |
Character or |
Value
Character label (e.g., "Cox HR", "OR",
"mean difference").
Estimate Propensity Scores
Description
Computes P(W = 1 \mid X) using the requested method, then
constructs both stabilized IPTW weights and the Bang & Robins (2005)
inverse propensity score (IPS) covariate. The data frame is returned
with three new columns: ps_hat (raw PS), sw (stabilized
IPTW weights), and ips_covar (IPS covariate).
Usage
estimate_propensity_scores(
data,
treat.name,
confounders.name,
method = c("grf", "lasso", "logistic", "none"),
grf_forest = NULL,
seed = 8316951L,
trim = c(0.025, 0.975)
)
Arguments
data |
Data frame. |
treat.name |
Character. Name of binary treatment column (0/1). |
confounders.name |
Character vector. Covariate names for the propensity model. |
method |
Character. One of |
grf_forest |
Optional. A fitted |
seed |
Integer. Random seed for GRF. |
trim |
Numeric vector of length 2. Quantile bounds for PS trimming.
Default: |
Value
A list with:
data |
Data frame with |
ps_hat |
Numeric vector of estimated propensity scores. |
sw |
Numeric vector of stabilized IPTW weights. |
ips_covar |
Numeric vector of inverse PS covariates (Bang & Robins 2005). |
method |
Character. Method actually used. |
trimmed |
Integer. Number of observations trimmed. |
References
Bang H, Robins JM (2005). "Doubly Robust Estimation in Missing Data and Causal Inference Models." Biometrics 61(4): 962–973.
Hernan MA, Robins JM, Brumback B (2000). "Marginal Structural Models and Causal Inference in Epidemiology." Epidemiology 11(5): 550–560.
Examples
## Not run:
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
ps_result <- estimate_propensity_scores(
data = df,
treat.name = "hormon",
confounders.name = c("age", "meno", "size", "grade3", "nodes",
"pgr", "er"),
method = "grf"
)
summary(ps_result$ps_hat)
summary(ps_result$sw)
## End(Not run)
Evaluate a Single Factor Combination with Status Tracking
Description
Tests whether a specific combination meets all criteria and returns a status code indicating how far the evaluation progressed.
Usage
evaluate_combination_with_status(
covs.in,
yy,
dd,
tt,
zz,
n.min,
d0.min,
d1.min,
hr.threshold,
minp,
rmin,
kk,
estimator_fn = NULL,
df_clean = NULL,
outcome_type = NULL,
adjust_covariates = NULL,
disable_effect_floor = FALSE
)
Arguments
covs.in |
Numeric vector. Factor selection indicators. |
yy |
Numeric vector. Outcome values. |
dd |
Numeric vector. Event indicators. |
tt |
Numeric vector. Treatment indicators. |
zz |
Matrix. Factor indicators. |
n.min |
Integer. Minimum sample size. |
d0.min |
Integer. Minimum control events (binary/survival) or ignored (continuous). |
d1.min |
Integer. Minimum treatment events (binary/survival) or ignored (continuous). |
hr.threshold |
Numeric. Effect threshold, on the scale this function
compares on: the natural hazard-ratio scale on the survival path, and
the resolved comparison scale (link for OR/RR/IRR, identity for
RD/IRD/MD) on the GLM paths. See |
minp |
Numeric. Minimum prevalence. |
rmin |
Integer. Minimum size reduction. |
kk |
Integer. Combination index. |
adjust_covariates |
Character vector or |
Value
List with:
- status
Integer status code: 0 = failed variance check, 1 = passed variance, failed prevalence, 2 = passed prevalence, failed redundancy, 3 = passed redundancy, failed per-arm filter, 4 = passed per-arm filter, failed sample size, 5 = passed sample size, failed model fit, 6 = passed model fit, failed effect threshold, 7 = passed all criteria (success)
- result
Result row if successful, NULL otherwise
Evaluate a Comparison Expression Without eval(parse())
Description
Parses a string of the form "var op value" and evaluates it
directly against a data frame column using operator dispatch. Falls back
to column-name lookup for bare names.
Usage
evaluate_comparison(expr, df)
Arguments
expr |
Character. An expression like |
df |
Data frame whose columns are referenced by |
Details
Supported operators (matched longest-first to avoid partial-match
ambiguity): <=, >=, !=, ==, <,
>.
If no operator is found, expr is treated as a column name and
the result is df[[expr]] == 1.
The value on the right-hand side is coerced to numeric when possible, otherwise kept as character for string comparisons.
Value
Logical vector of length nrow(df).
Examples
## Not run:
df <- data.frame(er = c(-1, 0, 1, 2), size = c(10, 20, 30, 40))
evaluate_comparison("er <= 0", df)
# [1] TRUE TRUE FALSE FALSE
evaluate_comparison("size > 25", df)
# [1] FALSE FALSE TRUE TRUE
## End(Not run)
Evaluate Consistency (Two-Stage Algorithm)
Description
Evaluates a single subgroup for consistency using a two-stage approach: Stage 1 screens with fewer splits, Stage 2 uses sequential batched evaluation with early stopping for efficient evaluation.
Usage
evaluate_consistency_twostage(
m,
index.Z,
names.Z,
df,
found.hrs,
hr.consistency,
pconsistency.threshold,
pconsistency.digits = 2,
maxk,
confs_labels,
details = FALSE,
n.splits.screen = 30,
screen.threshold = NULL,
n.splits.max = 400,
batch.size = 20,
conf.level = 0.95,
min.valid.screen = 10,
estimator_fn = NULL,
consistency_threshold = NULL,
adjust_covariates = NULL,
consistency_method = "split",
glm_resample_spec = NULL
)
Arguments
m |
Integer. Index of subgroup to evaluate. |
index.Z |
data.table or matrix. Factor indicators for all subgroups. |
names.Z |
Character vector. Names of factor columns. |
df |
data.frame. Original data with Y, Event, Treat, id columns. |
found.hrs |
data.table. Subgroup hazard ratio results. |
hr.consistency |
Numeric. |
pconsistency.threshold |
Numeric. |
pconsistency.digits |
Integer. Number of decimal places to which the
consistency proportion is rounded before it is compared with
|
maxk |
Integer. Maximum number of factors in a subgroup. |
confs_labels |
Character vector. Labels for confounders. |
details |
Logical. Print progress details. |
n.splits.screen |
Integer. Number of splits for Stage 1 (default 30). |
screen.threshold |
Numeric. Screening threshold for Stage 1 (default auto-calculated). |
n.splits.max |
Integer. Maximum total splits (default 400). |
batch.size |
Integer. Splits per batch in Stage 2 (default 20). |
conf.level |
Numeric. Confidence level for early stopping (default 0.95). |
min.valid.screen |
Integer. Minimum valid splits in Stage 1 (default 10). |
estimator_fn |
Closure or |
consistency_threshold |
Numeric or |
adjust_covariates |
Character vector or |
consistency_method |
Character. |
glm_resample_spec |
List or |
Value
Named numeric vector with consistency results, or NULL if not met.
Examples
## Not run:
# evaluate_consistency_twostage() is called internally by forestsearch()
# when use_twostage = TRUE. Use forestsearch() as the entry point:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status",
use_twostage = TRUE)
## End(Not run)
Cache and validate cut expressions efficiently
Description
Evaluates all cut expressions once and caches results to avoid redundant evaluation. Much faster than evaluating repeatedly.
Usage
evaluate_cuts_once(confs, df, details = FALSE)
Arguments
confs |
Character vector of cut expressions. |
df |
Data frame to evaluate expressions against. |
details |
Logical. Print details during execution. |
Details
This replaces multiple eval(parse()) calls scattered throughout get_FSdata. By caching results, we avoid:
Repeated parsing of expressions
Repeated evaluation on dataframe
Redundant uniqueness checks
Value
List with:
evaluations: List of evaluated vectors (logical TRUE/FALSE) for each cut
is_valid: Logical vector indicating which cuts produced >1 unique value
has_error: Logical vector indicating which cuts failed to evaluate
Examples
df <- data.frame(age = c(40, 55, 70), size = c(20, 30, 25))
confs <- c("age <= 50", "size > 25")
evaluate_cuts_once(confs, df)
Evaluate Subgroup Consistency (Single-Stage)
Description
Evaluates a single candidate subgroup for treatment-effect consistency using
repeated random data splits (the standard, non-two-stage algorithm). For
each split the subgroup effect is scored and compared against the harm
threshold; the consistency rate is the proportion of splits meeting it.
Called internally by forestsearch (via
subgroup.consistency); evaluate_consistency_twostage
is the screened, early-stopping counterpart used when
use_twostage = TRUE.
Usage
evaluate_subgroup_consistency(
m,
index.Z,
names.Z,
df,
found.hrs,
n.splits,
hr.consistency,
pconsistency.threshold,
pconsistency.digits = 2,
maxk,
confs_labels,
details = FALSE,
estimator_fn = NULL,
consistency_threshold = NULL,
adjust_covariates = NULL,
consistency_method = "split",
glm_resample_spec = NULL,
skip_consistency = FALSE
)
Arguments
m |
Integer. Index of subgroup to evaluate. |
index.Z |
data.table or matrix. Factor indicators for all subgroups. |
names.Z |
Character vector. Names of factor columns. |
df |
data.frame. Original data with Y, Event, Treat, id columns. |
found.hrs |
data.table. Subgroup hazard ratio results. |
n.splits |
Integer. Number of random splits used to evaluate consistency. |
hr.consistency |
Numeric. |
pconsistency.threshold |
Numeric. |
pconsistency.digits |
Integer. Number of decimal places to which the
consistency proportion is rounded before it is compared with
|
maxk |
Integer. Maximum number of factors in a subgroup. |
confs_labels |
Character vector. Labels for confounders. |
details |
Logical. Print progress details. |
estimator_fn |
Closure or |
consistency_threshold |
Numeric or |
adjust_covariates |
Character vector or |
consistency_method |
Character. |
glm_resample_spec |
List or |
skip_consistency |
Logical. When |
Value
Named numeric vector with consistency results, or NULL if the
subgroup does not meet the criteria.
Examples
## Not run:
# evaluate_subgroup_consistency() is called internally by forestsearch().
# Use forestsearch() as the entry point:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
## End(Not run)
Expand a seed covariate frame to n_super rows via a Gaussian copula.
Description
Expand a seed covariate frame to n_super rows via a Gaussian copula.
Usage
expand_covariates(
seed_df,
vars,
n_super,
seed = 20260903L,
integer_vars = NULL,
method = c("copula", "bootstrap")
)
Arguments
seed_df |
data.frame of the seed trial (only |
vars |
covariates to expand (numeric or 0/1 binary) |
n_super |
rows to generate |
seed |
RNG seed |
integer_vars |
variables to round back to integers (auto-detected if NULL) |
method |
"copula" (default) or "bootstrap" (row resampling, as the package does by default – no novel covariate patterns) |
Value
data.frame with n_super rows and columns vars
Describe the Pareto-Frontier Selection in Words
Description
Wraps pareto_frontier_table with a verbal account of
the selection: why the chosen subgroup satisfies the configured
sg_focus + selection_rule criterion, and, for each
non-dominated alternative, on which lexicographic axis it lost to
the winner.
Usage
explain_pareto_selection(
fs,
ci_table = NULL,
effect_label = NULL,
digits = 3L,
verbosity = c("medium", "compact", "detailed"),
format = c("markdown", "character", "list")
)
Arguments
fs |
A |
ci_table |
Optional data.table from
|
effect_label |
Optional character override for the effect-
measure label in prose, e.g. |
digits |
Integer. Decimal places for effect estimates and
bound values in the prose. Default |
verbosity |
One of |
format |
One of |
Details
The wrapper is descriptive, not algorithmic: it reads the
frontier and the configuration off the returned fs object and
renders prose. It does not re-run the selection. A future change
to the lexicographic key in forestsearch would require
updating this wrapper to remain accurate.
Value
Depending on format:
"markdown"A single string with embedded line breaks and bullet markers, ready for
cat()in a Quarto chunk withresults = "asis"."character"A character vector with one element per logical paragraph or bullet (no markdown markers).
"list"A structured list with named elements
criterion,selected,losers,closing, andmeta.
Verbosity
"compact"Criterion sentence + selected line + a one-line bullet per loser.
"medium"(default)Criterion sentence + selected paragraph (including the Naive CI if
ci_tableis given) + one short paragraph per loser, each naming the lex-key axis on which the loser was beaten."detailed"Medium + a closing paragraph that reproduces the formal lex key, the resolved indicator values, and a note about any ties broken downstream.
Loser ordering
Losers are sorted by "closeness-to-winning", which is operationalised
as: later loss-axis comes first, then smaller gap on
that axis. Specifically, a loser that lost on K (the last
tiebreaker) is reported before one that lost on N (an early
axis), and a loser that beat the winner on every axis until the
final one is reported before a loser that fell off the band entirely.
This puts the boundary cases first, which is the most informative
order for understanding how the rule made its choice.
See Also
pareto_frontier_table,
plot_pareto_frontier,
compute_frontier_cis.
Examples
## Not run:
ci_tab <- compute_frontier_cis(fs, n_splits = 1000, seed = 1)
# Inline in a Quarto chunk
cat(explain_pareto_selection(fs, ci_table = ci_tab))
# Or get the structured object
explain_pareto_selection(fs, ci_table = ci_tab, format = "list")
## End(Not run)
Extract all cuts from fitted trees
Description
Consolidates cut information from all fitted policy trees. This is the default behavior that returns cuts from all trees regardless of which tree identified the selected subgroup.
Usage
extract_all_tree_cuts(trees, maxdepth)
Arguments
trees |
List. Policy trees (indexed by depth) |
maxdepth |
Integer. Maximum tree depth |
Value
List with cuts and names for each tree and combined
Extract Subgroup Definition from ForestSearch Object
Description
Internal helper to extract human-readable subgroup definition.
Usage
extract_fs_subgroup_definition(fs.est, verbose = FALSE)
Arguments
fs.est |
A forestsearch object. |
verbose |
Logical. Print diagnostic messages. |
Value
Character string describing the subgroup definition.
Recover the GRF arguments used inside a forestsearch run
Description
Returns the exact list of arguments that forestsearch() passed
to its inner GRF subgroup-identification call
(grf.subg.harm.survival() for survival outcomes,
grf.subg.harm.glm() for binary, continuous, and count
outcomes). The result is suitable for
do.call(grf.subg.harm.glm, args) or
do.call(grf.subg.harm.survival, args) to reproduce the
standalone GRF run with identical settings.
Usage
extract_grf_args(fs)
Arguments
fs |
A fitted |
Details
Typical uses:
Sensitivity analyses: modify one or two arguments and re-fit GRF without re-running the full forestsearch pipeline, e.g.
do.call(grf.subg.harm.glm, modifyList(extract_grf_args(fs), list(maxdepth = 3L))).Reproducibility audits: confirm what the inner GRF saw.
Independent variable-importance or policy-tree extraction from a forest fit with the same configuration.
This function reads from fs$args_call_all (populated by
forestsearch() via mget() of its formals) and from
fs$outcome_type. It is the user-facing complement to the
internal builders that the forestsearch() call sites use.
Value
A named list of arguments. For survival outcomes the list
targets grf.subg.harm.survival; for binary,
continuous, or count outcomes it targets
grf.subg.harm.glm. The list carries an attribute
"grf_function" holding the target function name as a
character string.
See Also
forestsearch,
grf.subg.harm.glm,
grf.subg.harm.survival.
Examples
## Not run:
# Reproduce the inner GRF run
args <- extract_grf_args(fs)
fn <- match.fun(attr(args, "grf_function"))
grf_standalone <- do.call(fn, args)
# Sensitivity: deeper trees, everything else unchanged
args_deep <- modifyList(args, list(maxdepth = 3L))
grf_deep <- do.call(fn, args_deep)
## End(Not run)
Extract redundancy flag for subgroup combinations
Description
Checks if adding each factor to a subgroup reduces the sample size by at least rmin.
Usage
extract_idx_flagredundancy(x, rmin)
Arguments
x |
Matrix of subgroup factor indicators. |
rmin |
Integer. Minimum required reduction in sample size. |
Value
List with id.x (membership vector) and flag.redundant (logical).
Tidy summary of selected subgroups across a comparison run
Description
Reduces a forestsearch_comparison object to a tidy
data.frame with one row per combo, summarizing each configuration's
selected subgroup. Intended for post-hoc aggregation across
analyses (e.g.\ side-by-side summary of a continuous-outcome
comparison and a binary-outcome comparison in a single document).
Usage
extract_selected_subgroups(x, analysis_label = NULL)
Arguments
x |
A |
analysis_label |
Optional character scalar identifying the
parent analysis (e.g.\ |
Value
A data.frame with columns:
analysisIf
analysis_labelis supplied; otherwise omitted.comboInteger index,
1..n_combos.combo_labelPretty label (e.g.\
"effMaxSG / pareto").sg_focus,selection_ruleThe two comparison axes.
outcome_type,effect_measureRecorded per row from
fs$outcome_type/fs$effect_measure, so mixed-effect-measure runs remain comparable.subgroupSelected subgroup definition –
NA_character_if no subgroup was identified; otherwisepaste(fs$sg.harm, collapse = " & ").N_H,N_HcSample sizes in the harm subgroup H (
treat.recommend == 0) and its complement Hc.effectThe winner's effect estimate on the natural scale (
\expapplied to the storedhrcolumn for ratio-scale measuresOR/RR/IRR; passed through otherwise).PconsThe winner's consistency probability.
KNumber of factors in the winner's subgroup definition.
n_passedNumber of candidates that passed consistency (=
nrow(out_sg$result)).on_frontierLogical – did the selected subgroup lie on the Pareto frontier?
NAwhen the frontier is unavailable.errorThe error message for failed combos (
NA_character_on success).
Examples
## Not run:
cmp_cont <- compare_selection_rules(...) # continuous analysis
cmp_bin <- compare_selection_rules(...) # binary analysis
tidy_cont <- extract_selected_subgroups(cmp_cont, analysis_label = "continuous")
tidy_bin <- extract_selected_subgroups(cmp_bin, analysis_label = "binary")
all_selected <- rbind(tidy_cont, tidy_bin)
saveRDS(all_selected, "selected_subgroups_all.rds")
## End(Not run)
Extract cuts from selected tree only
Description
Extracts cut information only from the tree at the specified selected depth.
This provides a focused set of cuts from the tree that identified the
subgroup meeting the dmin.grf criterion, rather than cuts from all trees.
Usage
extract_selected_tree_cuts(trees, selected_depth, maxdepth)
Arguments
trees |
List. Policy trees (indexed by depth) |
selected_depth |
Integer. Depth of the selected tree (from best_subgroup$depth) |
maxdepth |
Integer. Maximum tree depth (for populating tree-specific slots) |
Details
This function is used when return_selected_cuts_only = TRUE in
grf.subg.harm.survival(). It returns:
-
tree1,tree2,tree3: Individual tree cuts (still populated for reference) -
names1,names2,names3: Individual tree variable names -
all: Cuts from the SELECTED tree only (not union of all trees) -
all_names: Variable names from the SELECTED tree only -
selected_depth: The depth that was selected
Value
List with cuts and names, structured similarly to extract_all_tree_cuts
but with only the selected tree's cuts in the all field
Examples
## Not run:
# When best_subgroup$depth = 2, only tree2 cuts are returned in $all
tree_cuts <- extract_selected_tree_cuts(trees, selected_depth = 2, maxdepth = 2)
tree_cuts$all
# Returns only cuts from depth 2 tree
## End(Not run)
Extract Subgroup Information
Description
Extracts subgroup definition and membership from results.
Usage
extract_subgroup(df, top_result, index.Z, names.Z, confs_labels)
Arguments
df |
Data.frame. Original analysis data. |
top_result |
Data.table row. Top subgroup result. |
index.Z |
Matrix. Factor indicators for all subgroups. |
names.Z |
Character vector. Factor column names. |
confs_labels |
Character vector. Human-readable labels. |
Value
List with sg.harm, sg.harm_label, df_flag, sg.harm.id.
Extract cut information from a policy tree
Description
Extracts all split points and variables from a policy tree
Usage
extract_tree_cuts(tree)
Arguments
tree |
Policy tree object |
Value
List with cuts (expressions) and names (unique variables)
Generate Figure Note for Quarto/RMarkdown
Description
Formats the figure note from plot_sg_weighted_km() output for use in Quarto or RMarkdown documents.
Usage
figure_note(
x,
prefix = "*Note*: ",
include_definition = TRUE,
include_hr_explanation = TRUE,
custom_text = NULL
)
Arguments
x |
Output from plot_sg_weighted_km() |
prefix |
Character. Prefix for the note. Default uses italic Note. |
include_definition |
Logical. Include subgroup definition. Default: TRUE |
include_hr_explanation |
Logical. Include HR(bc) explanation. Default: TRUE |
custom_text |
Character. Additional custom text to append. Default: NULL |
Value
Character string formatted as a figure note, or NULL if no content
Examples
## Not run:
km_result <- plot_sg_weighted_km(fs.est = fs)
cat(figure_note(km_result))
## End(Not run)
Filter a vector by LASSO-selected variables
Description
Returns elements of x that are in lassokeep.
Usage
filter_by_lassokeep(x, lassokeep)
Arguments
x |
Character vector. |
lassokeep |
Character vector of selected variables. |
Value
Filtered character vector or NULL.
Examples
filter_by_lassokeep(c("age", "size", "nodes"), lassokeep = c("age", "nodes"))
filter_by_lassokeep(c("pgr", "er"), lassokeep = c("age", "nodes")) # returns NULL
Filter and merge arguments for function calls
Description
Simplifies the common pattern of filtering arguments from a source list to match a target function's formal parameters, then adding/overriding specific arguments.
Usage
filter_call_args(source_args, target_func, override_args = NULL)
Arguments
source_args |
List of all arguments (typically from |
target_func |
Function whose formals define the filter criteria. |
override_args |
List of arguments to add or override (optional). |
Details
This function:
Extracts formal parameter names from
target_funcKeeps only arguments from
source_argsthat match those namesAdds or overrides with any
override_argsprovided
Reduces boilerplate and improves readability across the codebase.
Value
List of filtered arguments ready for do.call().
Examples
## Not run:
# Instead of:
args_FS <- names(formals(get_FSdata))
args_FS_filtered <- args_call_all[names(args_call_all) %in% args_FS]
args_FS_filtered$df.analysis <- df.analysis
args_FS_filtered$grf_cuts <- grf_cuts
FSdata <- do.call(get_FSdata, args_FS_filtered)
# You now write:
FSdata <- do.call(get_FSdata,
filter_call_args(args_call_all, get_FSdata,
list(df.analysis = df.analysis, grf_cuts = grf_cuts)))
## End(Not run)
Filter a Parameter List to a Function's Formals
Description
Given a named list of user-supplied parameters and a target function,
returns only the parameter entries whose names appear in the target's
formals. Entries with names that do not match any
formal (typos, renamed parameters, parameters from a different function)
are either silently dropped (default) or reported via warning().
Usage
filter_valid_args(
params,
target_fun,
exclude = character(0),
warn_unknown = FALSE
)
Arguments
params |
Named list of user-supplied parameter values. |
target_fun |
A function (not its name). Its |
exclude |
Character vector of formal-argument names to explicitly
exclude from the allowlist, even though they appear in
|
warn_unknown |
Logical. If |
Details
This replaces the older "hand-maintained valid_pnames whitelist"
pattern in .run_fs_analysis_gen() and .run_grf_analysis_gen().
New forestsearch() or grf.subg.harm.*() arguments flow
through automatically; authors of new parameters no longer need to
update a whitelist in a separate file to make the new parameter
reachable from user-level fs_params / grf_params lists.
Value
A named list, the subset of params whose names are in
formals(target_fun) and not in exclude.
Find Covariate Any Match
Description
Helper function to determine if any CV fold found a subgroup involving the same covariate (not necessarily same cut).
Usage
find_covariate_any_match(sg_target, sg1, sg2, confs)
Arguments
sg_target |
Character. Target subgroup definition to match. |
sg1 |
Character vector. Subgroup 1 labels for each fold. |
sg2 |
Character vector. Subgroup 2 labels for each fold. |
confs |
Character vector. Confounder names. |
Value
Numeric vector (0/1) indicating match for each fold.
Find k_inter Value to Achieve Target Harm Subgroup Hazard Ratio
Description
Uses numerical root-finding to determine the interaction parameter (k_inter) that achieves a specified target hazard ratio in the harm subgroup. This is the most efficient method for single target calibration.
Usage
find_k_inter_for_target_hr(
target_hr_harm,
data,
continuous_vars,
factor_vars,
outcome_var,
event_var,
treatment_var,
subgroup_vars,
subgroup_cuts,
k_treat = 1,
k_inter_range = c(-10, 10),
tol = 0.001,
tol_rel = 2.5,
n_super = 5000,
verbose = TRUE
)
Arguments
target_hr_harm |
Numeric value specifying the target hazard ratio for the harm subgroup. Must be positive. |
data |
A data.frame containing the dataset to use for model fitting. |
continuous_vars |
Character vector of continuous variable names to be standardized and included as covariates. |
factor_vars |
Character vector of factor/categorical variable names to be converted to dummy variables. |
outcome_var |
Character string specifying the name of the outcome/time variable. |
event_var |
Character string specifying the name of the event/status variable (1 = event, 0 = censored). |
treatment_var |
Character string specifying the name of the treatment variable. |
subgroup_vars |
Character vector of variable names defining the subgroup. |
subgroup_cuts |
Named list of cutpoint specifications for subgroup variables.
See |
k_treat |
Numeric value for treatment effect modifier. Default is 1 (no modification). |
k_inter_range |
Numeric vector of length 2 specifying the search range for k_inter. Default is c(-10, 10). |
tol |
Numeric value specifying tolerance for root finding convergence. Default is 0.001. |
tol_rel |
Numeric. Maximum tolerated relative error, in percent; the
call stops if the achieved HR is not within |
n_super |
Integer specifying size of super population for hazard ratio calculation. Default is 5000. |
verbose |
Logical indicating whether to print progress information. Default is TRUE. |
Details
This function uses the uniroot algorithm to solve the equation:
HR_{harm}(k_{inter}) - HR_{target} = 0
The algorithm typically converges within 5-10 iterations and achieves high precision (within the specified tolerance). If the root-finding fails, the function evaluates the boundaries and provides diagnostic information.
Value
A list of class "k_inter_result" containing:
- k_inter
Numeric value of optimal k_inter parameter
- achieved_hr_harm
Numeric value of achieved hazard ratio in harm subgroup
- target_hr_harm
Numeric value of target hazard ratio (for reference)
- error
Numeric value of absolute error between achieved and target HR
- dgm
Object of class "aft_dgm_flex" containing the final DGM
- convergence
Integer number of iterations to convergence
- method
Character string "root-finding" indicating method used
See Also
sensitivity_analysis_k_inter for sensitivity analysis
generate_aft_dgm_flex for DGM generation
Examples
## Not run:
gbsg <- survival::gbsg
# Find k_inter for target HR = 2.0 in harm subgroup
result <- find_k_inter_for_target_hr(
target_hr_harm = 2.0,
data = gbsg,
continuous_vars = c("age", "er", "pgr"),
factor_vars = c("meno", "grade"),
outcome_var = "rfstime",
event_var = "status",
treatment_var = "hormon",
subgroup_vars = c("er", "meno"),
subgroup_cuts = list(
er = list(type = "quantile", value = 0.25),
meno = 0
),
k_treat = 1.0,
verbose = TRUE
)
cat("Optimal k_inter:", result$k_inter, "\n")
cat("Achieved HR:", result$achieved_hr_harm, "\n")
## End(Not run)
Find the split that leads to a specific leaf node (deprecated)
Description
Deprecated.
Returns only the single split immediately above leaf_node, rendered as
"var <= value". This is not a correct subgroup definition for a leaf
below depth 1: it drops the rest of the root-to-leaf path and mis-renders
right-turns (which are >, not <=). Use the path-based
.grf_build_subgroup_definition() instead, which
grf.subg.harm.survival() now calls internally. Retained (exported) only
for backward compatibility; it is no longer used inside the package and
will be removed in a future release.
Usage
find_leaf_split(tree, leaf_node)
Arguments
tree |
Policy tree object |
leaf_node |
Integer. Leaf node identifier |
Value
Character string with a single split expression, or NULL.
See Also
grf.subg.harm.survival for the standard entry point.
Examples
## Not run:
# Deprecated; grf.subg.harm.survival() builds the full path internally.
## End(Not run)
Find Quantile for Target Subgroup Proportion
Description
Determines the quantile cutpoint that achieves a target proportion of observations in a subgroup. Useful for calibrating subgroup sizes.
Usage
find_quantile_for_proportion(
data,
var_name,
target_prop,
direction = "less",
tol = 1e-04
)
Arguments
data |
A data.frame containing the variable of interest |
var_name |
Character string specifying the variable name to analyze |
target_prop |
Numeric value between 0 and 1 specifying the target proportion of observations to be included in the subgroup |
direction |
Character string: "less" for values <= cutpoint (default), "greater" for values > cutpoint |
tol |
Numeric tolerance for root finding algorithm. Default is 0.0001 |
Details
This function uses root finding (uniroot) to determine the quantile
that results in exactly the target proportion of observations being classified
into the subgroup. This is particularly useful when you want to ensure a
specific subgroup size regardless of the data distribution.
Value
A list containing:
- quantile
The quantile value (between 0 and 1) that achieves the target proportion
- cutpoint
The actual data value corresponding to this quantile
- actual_proportion
The achieved proportion (should equal target_prop within tolerance)
See Also
Examples
## Not run:
gbsg <- survival::gbsg
# Find ER cutpoint for 12.5% subgroup
result <- find_quantile_for_proportion(
data = gbsg,
var_name = "er",
target_prop = 0.125,
direction = "less"
)
print(result)
# Use in subgroup definition
subgroup_cuts = list(
er = list(type = "quantile", value = result$quantile)
)
## End(Not run)
Find Minimum Sample Size for Target Detection Power
Description
Determines the minimum subgroup sample size needed to achieve a target detection probability for a given true hazard ratio.
Usage
find_required_sample_size(
theta,
target_power = 0.8,
prop_cens = 0.3,
hr_threshold = 1.25,
hr_consistency = 1,
n_range = c(20L, 500L),
tol = 1,
verbose = TRUE
)
Arguments
theta |
Numeric. True hazard ratio in subgroup. |
target_power |
Numeric. Target detection probability (0-1). Default: 0.80 |
prop_cens |
Numeric. Proportion censored. Default: 0.3 |
hr_threshold |
Numeric. HR threshold. Default: 1.25 |
hr_consistency |
Numeric. HR consistency threshold. Default: 1.0 |
n_range |
Integer vector of length 2. Range of sample sizes to search. Default: c(20, 500) |
tol |
Numeric. Tolerance for bisection search. Default: 1 |
verbose |
Logical. Print progress. Default: TRUE |
Value
A list with:
n_sg_required |
Minimum sample size (rounded up) |
achieved_power |
Actual detection probability at n_sg_required |
theta |
Input hazard ratio |
target_power |
Input target power |
Examples
## Not run:
# Find sample size for 80% power to detect HR = 1.5
result <- find_required_sample_size(
theta = 1.5,
target_power = 0.80,
prop_cens = 0.2
)
print(result)
## End(Not run)
Fit AFT Model with Optional Spline Treatment Effect
Description
Fit AFT Model with Optional Spline Treatment Effect
Usage
fit_aft_model(
df_work,
interaction_term,
k_treat,
k_inter,
verbose,
spline_spec = NULL,
set_var = NULL,
beta_var = NULL
)
Fit AFT Model and Apply Effect Modifiers
Description
Fit AFT Model and Apply Effect Modifiers
Usage
fit_aft_model_legacy(df_work, interaction_term, k_treat, k_inter, verbose)
Fit AFT Model with Spline Treatment Effect
Description
Fit AFT Model with Spline Treatment Effect
Usage
fit_aft_model_spline(
df_work,
covariate_cols,
interaction_term,
k_treat,
k_inter,
spline_spec,
verbose,
set_var = NULL,
beta_var = NULL
)
Fit Standard AFT Model (Non-Spline)
Description
Fit Standard AFT Model (Non-Spline)
Usage
fit_aft_model_standard(
df_work,
covariate_cols,
interaction_term,
k_treat,
k_inter,
verbose,
set_var = NULL,
beta_var = NULL
)
Fit causal survival forest
Description
Wrapper function to fit GRF causal survival forest with appropriate settings
Usage
fit_causal_forest(X, Y, W, D, tau.rmst, RCT, seedit, tune_grf = FALSE)
Arguments
X |
Matrix. Covariate matrix |
Y |
Numeric vector. Outcome variable |
W |
Numeric vector. Treatment indicator |
D |
Numeric vector. Event indicator |
tau.rmst |
Numeric. Time horizon for RMST |
RCT |
Logical. Is this RCT data? |
seedit |
Integer. Random seed |
tune_grf |
Logical. If TRUE, enables cross-validated hyperparameter
tuning via |
Value
Causal survival forest object
Examples
library(survival)
library(grf)
df <- survival::gbsg
X <- as.matrix(df[, c("age", "meno", "size", "nodes", "pgr", "er")])
tau <- quantile(df$rfstime[df$status == 1], 0.6)
cs <- fit_causal_forest(X = X, Y = df$rfstime, W = df$hormon,
D = df$status, tau.rmst = tau,
RCT = FALSE, seedit = 42L)
Fit Causal Forest for GLM Outcomes
Description
Fits a GRF causal forest for binary or continuous outcomes using
grf::causal_forest(). This is the GLM counterpart to
fit_causal_forest, which uses causal_survival_forest.
Usage
fit_causal_forest_glm(X, Y, W, RCT, seedit, tune_grf = FALSE)
Arguments
X |
Numeric matrix. Covariates. |
Y |
Numeric vector. Outcome (binary 0/1 or continuous). |
W |
Numeric vector. Treatment indicator (0/1). |
RCT |
Logical. Is this RCT data? |
seedit |
Integer. Random seed. |
tune_grf |
Logical. If TRUE, enables cross-validated hyperparameter
tuning via |
Value
Causal forest object.
Fit Cox Model for Subgroup
Description
Fit Cox Model for Subgroup
Usage
fit_cox_for_subgroup(
yy,
dd,
tt,
id.x,
df_clean = NULL,
adjust_covariates = NULL
)
Arguments
yy, dd, tt, id.x |
Numeric vectors: outcome, event, treatment, and the subgroup membership indicator (1 = in subgroup). |
df_clean |
Data frame or |
adjust_covariates |
Character vector or |
Fit Cox Models for Subgroups
Description
Fits Cox models for two subgroups defined by treatment recommendation.
Usage
fit_cox_models(df, formula, treat.name = NULL)
Arguments
df |
Data frame. |
formula |
Cox model formula. |
treat.name |
Character or |
Value
List with HR and SE for each subgroup.
Examples
## Not run:
library(survival)
df <- data.frame(
tte = gbsg$rfstime / 30.4375,
event = gbsg$status,
treat = gbsg$hormon,
treat.recommend = as.integer(gbsg$er > 0)
)
formula <- build_cox_formula("tte", "event", "treat")
fit_cox_models(df, formula)
## End(Not run)
Fit Effect Models for Both Subgroups (GLM path)
Description
GLM counterpart to fit_cox_models. Calls the estimator closure on
H (treat.recommend == 0) and Hc (treat.recommend == 1) subgroups.
Usage
fit_effect_models(df, estimator_fn)
Arguments
df |
Data frame with a |
estimator_fn |
Closure from |
Value
List with H_obs, seH_obs, Hc_obs, seHc_obs.
Fit GLM for Subgroup via Estimator Closure
Description
Subsets the analysis data frame to the candidate subgroup and calls the
pre-built estimator closure. Returns a list compatible with
create_result_row so the downstream pipeline is unchanged.
Usage
fit_glm_for_subgroup(df_clean, id.x, estimator_fn)
Arguments
df_clean |
Data frame. The analysis data, row-aligned with the cleaned vectors. |
id.x |
Integer vector. 1 = subject in this subgroup, 0 = not. |
estimator_fn |
Closure from |
Details
For binary outcomes with effect_measure = "RD", the hr slot contains
the risk difference (not a hazard ratio) – the name is retained for
pipeline compatibility. Similarly, lower/upper are Wald-type CI
bounds and med0/med1 are NA.
Value
A list with components hr, lower, upper, med0, med1,
or NULL on failure.
Fit policy trees up to specified depth
Description
Fits policy trees of depths 1 through maxdepth and computes metrics
Usage
fit_policy_trees(X, data, dr.scores, maxdepth, n.min)
Arguments
X |
Matrix. Covariate matrix |
data |
Data frame. Original data |
dr.scores |
Matrix. Doubly robust scores |
maxdepth |
Integer. Maximum tree depth (1-3) |
n.min |
Integer. Minimum subgroup size |
Value
List with trees and combined values
Examples
## Not run:
# fit_policy_trees() is called internally by grf.subg.harm.survival().
# See grf.subg.harm.survival() for the standard entry point.
## End(Not run)
Estimate Subgroup Effect via Estimator Closure
Description
GLM counterpart to get_Cox_sg. Calls the pre-built estimator closure
on a data slice and returns est_obs and se_obs in the same format
that get_Cox_sg() returns.
Usage
fit_subgroup_effect(df_sg, estimator_fn)
Arguments
df_sg |
Data frame. The subgroup data slice. |
estimator_fn |
Closure from |
Value
List with est_obs (effect estimate) and se_obs (standard error).
Recommended figure height for a forest panel
Description
Converts a panel's row count to the fig.height (inches) used by the
vignette chunks: rows * row_height + overhead, rounded to 0.1.
Usage
forest_height(
x,
panel = c("single", "combo", "highrisk"),
row_height = 0.45,
overhead = 1.5
)
Arguments
x |
A |
panel |
|
row_height |
Inches per row (default 0.45). |
overhead |
Fixed inches for title, x-axis, and footnote (default 1.5). |
Value
A single numeric, rounded to one decimal.
See Also
Examples
## Not run:
fh_single <- forest_height(S, "single")
## End(Not run)
ForestSearch: Exploratory Subgroup Identification
Description
Implements advanced statistical methods for exploratory subgroup identification in clinical trials. Provides tools for identifying patient subgroups with differential treatment effects using machine learning approaches including Generalized Random Forests ('GRF'), LASSO regularization, and exhaustive combinatorial search algorithms. Supports survival endpoints (Cox proportional hazards), binary outcomes (log odds ratio, log relative risk, risk difference), continuous outcomes (mean difference), and count / rate outcomes (log incidence rate ratio via Poisson, quasi-Poisson, or negative-binomial GLMs with optional person-time offset). Features bootstrap bias correction using infinitesimal jackknife methods to address selection bias in post-hoc analyses. Designed for clinical researchers conducting exploratory subgroup analyses in randomized controlled trials, particularly for multi-regional clinical trials ('MRCT') requiring regional consistency evaluation. Methods are described in Leon et al. (2024) doi:10.1002/sim.10163.
Usage
forestsearch(
df.analysis,
outcome.name = "tte",
event.name = "event",
treat.name = "treat",
id.name = "id",
potentialOutcome.name = NULL,
flag_harm.name = NULL,
confounders.name = NULL,
parallel_args = list(plan = "multisession", workers = .default_parallel_workers(),
show_message = TRUE),
df.predict = NULL,
df.test = NULL,
is.RCT = TRUE,
seedit = 8316951,
est.scale = "hr",
use_lasso = FALSE,
use_grf = FALSE,
grf_res = NULL,
grf_cuts = NULL,
use_dina = FALSE,
dina_res = NULL,
dina_cuts = NULL,
dina_args = list(),
dina_select_statistic = c("effect", "dina"),
subgroup_method = c("consistency", "dina", "grf"),
max_n_confounders = 1000,
grf_depth = 2,
grf_selection = c("frontier", "tree"),
grf_select_statistic = c("effect", "dr"),
dmin.grf = 0,
frac.tau = 0.8,
return_selected_cuts_only = TRUE,
conf_force = NULL,
defaultcut_names = NULL,
cut_type = "default",
exclude_cuts = NULL,
replace_med_grf = FALSE,
cont.cutoff = 4,
conf.cont_medians = NULL,
conf.cont_medians_force = NULL,
conf.cont_jcuts = NULL,
collapse_cuts = TRUE,
collapse_cuts_args = list(),
n.min = 60,
n.min.frac = 0.1,
effect.threshold = NULL,
consistency.threshold = NULL,
hr.threshold = 1.25,
hr.consistency = 1,
sg_focus = "hr",
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
fs.splits = 1000,
m1.threshold = Inf,
pconsistency.threshold = 0.9,
stop_threshold = pconsistency.threshold,
pconsistency.digits = 2,
show_candidate_summary = FALSE,
max_print = 10,
d0.min = 10,
d1.min = 10,
max.minutes = 3,
minp = 0.025,
details = FALSE,
quiet = FALSE,
maxk = 2,
by.risk = 12,
plot.sg = FALSE,
plot.grf = FALSE,
max_subgroups_search = Inf,
vi.grf.min = NULL,
use_twostage = TRUE,
twostage_args = list(),
consistency_method = c("resample", "split"),
outcome_type = c("survival", "binary", "continuous", "count"),
effect_measure = NULL,
offset.name = NULL,
adverse_outcome = NULL,
overdispersion = c("none", "quasi", "negbin"),
grf_count_transform = c("log", "identity"),
tune_grf = FALSE,
adjust_covariates = NULL,
ps_method = NULL,
ps_adjust_method = c("none", "iptw", "dr_gcomp"),
ps_hat = NULL,
mr_inference = FALSE,
mr_inference_args = list()
)
Arguments
df.analysis |
Data frame. Analysis dataset with required columns. | ||||||||||||||||||||||||
outcome.name |
Character. Name of time-to-event outcome variable. Default "tte". | ||||||||||||||||||||||||
event.name |
Character. Name of event indicator (1=event, 0=censored). Default "event". | ||||||||||||||||||||||||
treat.name |
Character. Name of treatment variable (1=treatment, 0=control). Default "treat". | ||||||||||||||||||||||||
id.name |
Character. Name of subject ID variable. Default "id". | ||||||||||||||||||||||||
potentialOutcome.name |
Character. Name of potential outcome variable (optional). | ||||||||||||||||||||||||
flag_harm.name |
Character. Name of true harm flag for simulations (optional). | ||||||||||||||||||||||||
confounders.name |
Character vector. Names of candidate subgroup-defining variables. | ||||||||||||||||||||||||
parallel_args |
List. Parallel processing configuration:
To disable parallelism entirely, pass
| ||||||||||||||||||||||||
df.predict |
Data frame. Prediction dataset (optional). | ||||||||||||||||||||||||
df.test |
Data frame. Test dataset (optional). | ||||||||||||||||||||||||
is.RCT |
Logical. Is this a randomized controlled trial? Default TRUE. | ||||||||||||||||||||||||
seedit |
Integer. Random seed. Default 8316951. | ||||||||||||||||||||||||
est.scale |
Character. Estimation scale ("hr" or "rmst"). Default "hr". | ||||||||||||||||||||||||
use_lasso |
Logical. Use LASSO for variable selection. Default FALSE.
The default changed from | ||||||||||||||||||||||||
use_grf |
Logical. Generate additional candidate cuts from a GRF
fit: tree-derived cutpoints are extracted via
| ||||||||||||||||||||||||
grf_res |
GRF results object (optional, for reuse). | ||||||||||||||||||||||||
grf_cuts |
List. Custom GRF cut points (optional). | ||||||||||||||||||||||||
use_dina |
Logical. Generate additional screening-stage candidate
cuts from a DINA per-covariate Pareto frontier, fed into
| ||||||||||||||||||||||||
dina_res |
A fitted DINA object (optional, for reuse). When NULL
and | ||||||||||||||||||||||||
dina_cuts |
Character vector of pre-supplied DINA cut expressions (optional). When supplied, the DINA fit / frontier step is skipped and these are used directly. | ||||||||||||||||||||||||
dina_args |
Named list of DINA tuning options, all optional.
Recognised keys:
The same holds for every frontier key, not only
Under | ||||||||||||||||||||||||
dina_select_statistic |
Character, one of | ||||||||||||||||||||||||
subgroup_method |
Character, one of | ||||||||||||||||||||||||
max_n_confounders |
Integer. Maximum number of candidate cut columns
retained after GRF variable-importance screening, keeping the
highest-importance columns. Default 1000. Applied only when
| ||||||||||||||||||||||||
grf_depth |
Integer. GRF tree depth. Default 2. | ||||||||||||||||||||||||
grf_selection |
Character, one of | ||||||||||||||||||||||||
grf_select_statistic |
Character, one of | ||||||||||||||||||||||||
dmin.grf |
Numeric. Minimum events for GRF. Default 0.0. | ||||||||||||||||||||||||
frac.tau |
Numeric in (0, 1]. Multiplier on the GRF time horizon
passed to | ||||||||||||||||||||||||
return_selected_cuts_only |
Logical. If TRUE (default), GRF returns only cuts from the
tree depth that identified the selected subgroup meeting | ||||||||||||||||||||||||
conf_force |
Character vector of forced cut expressions (optional),
e.g. | ||||||||||||||||||||||||
defaultcut_names |
Character vector. Default cut variable names (optional). | ||||||||||||||||||||||||
cut_type |
Character. Cut strategy: | ||||||||||||||||||||||||
exclude_cuts |
Character vector. Variables to exclude from cutting (optional). | ||||||||||||||||||||||||
replace_med_grf |
Logical. Replace median with GRF cuts. Default FALSE. | ||||||||||||||||||||||||
cont.cutoff |
Integer. Cutoff for continuous vs categorical. Default 4. | ||||||||||||||||||||||||
conf.cont_medians |
Character vector of continuous confounders to cut at
their median (optional). Names, not values – the cut location is the
median computed from the supplied data, so it is not fixed by this
argument and moves with a resample. To pin a cut location, pass a literal
cut expression via | ||||||||||||||||||||||||
conf.cont_medians_force |
Character vector of additional continuous
confounders to force a median cut on (optional). Names, not values, with
the same consequence as | ||||||||||||||||||||||||
conf.cont_jcuts |
Named integer list (each value >= 1). Per-variable
J-quantile cut override for continuous covariates: for an entry
| ||||||||||||||||||||||||
collapse_cuts |
Logical. If TRUE, collapse near-redundant continuous
candidate cuts before the search. Continuous covariates can generate many
thresholds that are practically redundant (e.g. | ||||||||||||||||||||||||
collapse_cuts_args |
List of overrides merged onto the defaults
| ||||||||||||||||||||||||
n.min |
Integer or | ||||||||||||||||||||||||
n.min.frac |
Numeric in (0, 1). Fraction of the analysis sample size
used for the adaptive | ||||||||||||||||||||||||
effect.threshold |
Numeric or NULL. | ||||||||||||||||||||||||
consistency.threshold |
Numeric or NULL. | ||||||||||||||||||||||||
hr.threshold |
Numeric. Legacy name for | ||||||||||||||||||||||||
hr.consistency |
Numeric. Legacy name for
| ||||||||||||||||||||||||
sg_focus |
Character. Subgroup selection focus – what is
selected, not merely how candidates are sorted. Except for
The descriptions above are for For Per-engine resolution. Every accepted spelling, and the rule that actually runs for it on each engine:
The collapse sets, stated explicitly. On
Default Scope of the band arguments. Choosing between the three effect-oriented rules.
| ||||||||||||||||||||||||
selection_rule |
Character. Rule defining the candidate
inclusion set for | ||||||||||||||||||||||||
effect_neighborhood |
Numeric in | ||||||||||||||||||||||||
fs.splits |
Integer. Number of splits for consistency evaluation (or maximum
splits when | ||||||||||||||||||||||||
m1.threshold |
Numeric. Maximum median survival threshold. Default Inf. | ||||||||||||||||||||||||
pconsistency.threshold |
Numeric. | ||||||||||||||||||||||||
stop_threshold |
Numeric in Applies to For every other focus
Note: Values > 1.0 are not permitted. To disable early
stopping, use | ||||||||||||||||||||||||
pconsistency.digits |
Integer. Number of decimal places to which the
consistency proportion is rounded before it is compared with
| ||||||||||||||||||||||||
show_candidate_summary |
Logical. If | ||||||||||||||||||||||||
max_print |
Integer. Maximum number of candidate subgroups to print in
the pre- and post-consistency candidate summaries when
| ||||||||||||||||||||||||
d0.min |
Integer. Minimum per-arm filter for candidate subgroups.
For | ||||||||||||||||||||||||
d1.min |
Integer. Same as | ||||||||||||||||||||||||
max.minutes |
Numeric. Currently inert; scheduled for
deprecation in v0.3.0. Previously intended as a wall-clock time
budget for the combination search, this argument is no longer
enforced in the parallelized search path and has no effect on
behavior. Search scope is governed by | ||||||||||||||||||||||||
minp |
Numeric. Minimum prevalence threshold. Default 0.025. | ||||||||||||||||||||||||
details |
Logical. Print progress details. Default FALSE. | ||||||||||||||||||||||||
quiet |
Logical. If TRUE, suppress the configuration summary
message printed at startup. Useful when | ||||||||||||||||||||||||
maxk |
Integer. Maximum number of factors per subgroup. Default 2. | ||||||||||||||||||||||||
by.risk |
Integer. Risk table interval. Default 12. | ||||||||||||||||||||||||
plot.sg |
Logical. Plot subgroup survival curves. Default FALSE. | ||||||||||||||||||||||||
plot.grf |
Logical. Plot GRF results. Default FALSE. | ||||||||||||||||||||||||
max_subgroups_search |
Numeric. Maximum subgroups to evaluate.
Default Set a finite value only to bound runtime, knowing it can change which
subgroup is selected. The default was | ||||||||||||||||||||||||
vi.grf.min |
Numeric or | ||||||||||||||||||||||||
use_twostage |
Logical. Use two-stage sequential consistency algorithm for
improved performance. Default FALSE for backward compatibility. When TRUE,
| ||||||||||||||||||||||||
twostage_args |
List. Parameters for two-stage algorithm (only used when
| ||||||||||||||||||||||||
consistency_method |
Character. | ||||||||||||||||||||||||
outcome_type |
Character. One of | ||||||||||||||||||||||||
effect_measure |
Character or | ||||||||||||||||||||||||
offset.name |
Character or | ||||||||||||||||||||||||
adverse_outcome |
Logical or | ||||||||||||||||||||||||
overdispersion |
Character. Overdispersion correction for
| ||||||||||||||||||||||||
grf_count_transform |
Character. Transformation applied to the
count outcome before passing to | ||||||||||||||||||||||||
tune_grf |
Logical. If | ||||||||||||||||||||||||
adjust_covariates |
Character vector or | ||||||||||||||||||||||||
ps_method |
Character or GLM outcome types only ( | ||||||||||||||||||||||||
ps_adjust_method |
Character. PS adjustment method: GLM outcome types only, and consumed only when | ||||||||||||||||||||||||
ps_hat |
Numeric vector or Not carried into bootstrap replicates or cross-validation folds: a
resample has the same row count but different subjects, so a supplied
score would be attached to the wrong people. Both re-estimate their own
score per replicate/fold under | ||||||||||||||||||||||||
mr_inference |
Logical. Whether to follow the search with multiplier
resampling (MR): a refit-free de-biased estimate of the selected
subgroup's treatment effect, an interval for it, and a harm-confirmation
flag. Default | ||||||||||||||||||||||||
mr_inference_args |
List of optional MR controls; used only when
|
Details
Algorithm Overview:
-
Variable Selection: GRF identifies variables with treatment effect heterogeneity; LASSO selects most predictive
-
Subgroup Discovery: Exhaustive search over factor combinations up to
maxk -
Consistency Validation: Split-sample validation ensures reproducibility
-
Selection: Choose subgroup based on
sg_focuscriterion
Two-Stage Consistency Algorithm:
When use_twostage = TRUE, the consistency evaluation uses an optimized
algorithm that can provide 3-10x speedup:
-
Stage 1: Quick screening with
n.splits.screensplits eliminates clearly non-viable candidates -
Stage 2: Sequential batched evaluation with early stopping for candidates passing Stage 1
The two-stage algorithm is recommended for:
Exploratory analyses with many candidate subgroups
Large
fs.splitsvalues (>200)Iterative model development
For final regulatory submissions, use_twostage = FALSE may be preferred
for exact reproducibility.
Value
A list of class "forestsearch" containing:
- grp.consistency
Consistency evaluation results including:
out_sg: Selected subgroup based on sg_focus
sg_focus: Focus criterion used
df_flag: Treatment recommendations
sg.harm: Subgroup definition labels (character vector of cut names) – same as the top-level
sg.harmsg.harm.id: per-subject 0/1 membership indicator (not a character vector of cut expressions; see the Field naming collision section below)
algorithm: "twostage" or "fixed"
n_candidates_evaluated: Number evaluated
n_passed: Number passing threshold
- find.grps
Subgroup search results
- confounders.candidate
Candidate confounders considered
- confounders.evaluated
Confounders after variable selection
- df.est
Analysis data with treatment recommendations
- df.predict
Prediction data with recommendations (if provided)
- df.test
Test data with recommendations (if provided)
- minutes_all
Total computation time
- grf_res
GRF results object
- sg_focus
Subgroup focus criterion used
- sg.harm
Selected subgroup definition – character vector of factor-level cut names. Recommended accessor for displaying the identified subgroup (e.g.,
paste(obj$sg.harm, collapse = " & ")).NULLif no subgroup was identified.- grf_cuts
GRF cut points used
- prop_maxk
Proportion of max combinations searched
- max_sg_est
Maximum subgroup HR estimate
- grf_plot
GRF plot object (if plot.grf = TRUE)
- args_call_all
All arguments for reproducibility
- threshold_config
The resolved threshold configuration:
outcome_type,effect_measure,screening,consistency,screening_natural,consistency_natural,scale,pconsistencyand two description strings. Read$scalebefore reading$screening. The phrase "the screening threshold" names two different numbers depending on which object is inspected:On the survival path
$screeningislog(hr.threshold)– the scale the admission set and multiplier resampling work on – while the candidate search itself compares the fitted hazard ratio againsthr.thresholdon the natural scale. The two are the same threshold on two scales, and$screening_naturalholds the search's value.On the GLM paths
$screeningis the value the search compares against directly: the link scale for ratio measures (logof the supplied ratio) and the identity scale forRD,IRDandMD.
$consistencyfollows the same convention, and$pconsistencyis the ratepconsistency.threshold, which has no scale.- family_status
Character scalar recording whether the candidate family the identifier ranked over is fixed – the condition multiplier resampling requires (see
mr_inference)."no-front-end":subgroup_method = "consistency"withuse_lasso,use_grfanduse_dinaallFALSE. Named for what it checks: no fitted model shapes the family on the observed data. This is weaker than the manuscript's Section 2.1 fixed family, which additionally requires resample- invariant cut locations – quantile-derived cuts are not."conditional-removable": the same method with at least one front end on – data-dependent, but fixed by turning those off."conditional-inherent":subgroup_methodof"dina"or"grf", where the family is generated by fitting a model to the same data and cannot be made fixed. Descriptive only; it never raises a condition.- mr_inference
Multiplier-resampling (MR) result, or
NULLwhen MR did not run. Seefs_mr_inferencefor fields (naive,debiased,selection_bias,settings,harm_flag,timing_seconds). MR approximates the full bootstrap (FB) to leading order; its interval is the infinitesimal-jackknife analogue of the FB interval under the defaultci_method = "field"and under"ij", and the subgroup robust-SE interval under"wald". Under the default the return also carries thefieldelement.- mr_harm_confirmed
Logical, three-valued.
TRUEwhen MR ran and the de-biased estimate still indicates harm underconfirm_rule;FALSEwhen MR ran and it does not;NAwhen MR did not run or could not be computed.NAis therefore absence of evidence, not evidence against harm – test it withis.na()beforeisTRUE(), which collapsesNAtoFALSEand so reports "harm not confirmed" for an analysis that never ran. See the vocabulary section.
Threshold vocabulary and resolved defaults
Three arguments carry thresholds, and two of them have the word "consistency" in the name. They are different quantities:
c1–effect.threshold(legacyhr.threshold)The screening threshold on a candidate subgroup's own effect, estimated on the full sample. A candidate must clear it to enter the consistency stage.
c2–consistency.threshold(legacyhr.consistency)An effect threshold, applied to the subgroup's effect estimated within each split half. Not a proportion, and not the consistency rate.
p*–pconsistency.thresholdA proportion in
[0, 1]: the fraction of splits that must clearc2. Not an effect threshold.
In one line: c2 sets the bar a split half must clear, and p* counts how
often that bar is met.
Resolved defaults, per estimand. Reading
hr.consistency's default as "1.0" is misleading on the GLM paths:
when neither threshold is supplied, the value a run actually uses is
resolved from the estimand. The resolved pairs are:
| resolved estimand | c1 | c2 |
HR (survival) | 1.25, compared by the search on the natural HR scale | 1.0 |
OR, RR, IRR | log(1.25) | log(1.0), i.e. 0 |
RD | 0.05 | 0.0 |
IRD | 0.01 | 0.0 |
MD | 0.0 | 0.0 |
Supplying effect.threshold or consistency.threshold (or the
legacy names) suppresses the corresponding remap and the supplied value is
used on the scale shown above.
The pair rule, on the consistency path with a ratio estimand. For
subgroup_method = "consistency" with survival (HR) or binary
effect_measure = "OR" – the two estimands whose c1 and c2 sit on
one comparable ratio scale – the two thresholds are treated as a pair:
-
c2 > c1is an error. It is degenerate: the consistency-stage entry condition re-screens the admitted family onc2, so candidates the screen let through are discarded at the stage boundary, and when none survives the consistency stage does not run at all and the fit reads as a finding of no subgroup.c2 = c1is allowed. -
c1supplied withoutc2derivesc2 = 0.80 * c1on the ratio scale (on the log scale an additive shift oflog(0.80), never0.80 * log(c1)): 1.25 -> 1.00, 1.00 -> 0.80, 0.90 -> 0.72. Announced with a message, and carried into every bootstrap replicate and CV fold. The default pair (1.25, 1.0) is the rule's fixed point, so a call that sets neither threshold is unaffected. An explicitly suppliedc2, in either spelling, is never overridden. Supplying both spellings of one threshold at disagreeing values is an error naming both.
Neither rule applies under subgroup_method = "dina" or
"grf" (which return before the consistency stage and never consult
c2), nor to the identity-scale estimands RD, IRD,
MD, nor to IRR.
pconsistency.threshold is not remapped. It is a rate, and
it is identical for every outcome_type and every estimand.
Field naming collision with GRF results
The top-level sg.harm on this object and the
grp.consistency$sg.harm nested field both hold a character
vector of cut expressions. The nested grp.consistency$sg.harm.id
field holds a per-subject 0/1 membership indicator, not a cut
vector – unlike sg.harm.id on GRF result objects, which
does hold cut expressions. See
subgroup.consistency and grf.subg.harm.glm
for the full discussion of this naming collision. Code that must
handle both object types should prefer the top-level sg.harm
accessor on forestsearch results.
Vocabulary FB MR and harm confirmation
Four different things in this package used to share the word "gate". They are distinct, they happen at different stages, and only the last one is a pass/fail decision:
- Admissibility criteria
Which candidate subgroups are even eligible to be considered:
hr.threshold,pconsistency.threshold,n.min,d0.min,d1.min,maxk. Applied during the search.- Selection rule
Which admissible candidate is chosen as the estimated subgroup:
sg_focus. Applied during the search.- Post-selection inference
Correcting the chosen subgroup's treatment effect for the optimism of having chosen it, and putting an interval around the corrected value. This is FB and MR below. Applied after the search.
- Harm confirmation
The one pass/fail: does the corrected estimate still indicate harm? This is
confirm_ruleagainstt_confirm, reported asmr_harm_confirmed.
Full bootstrap (FB) — forestsearch_bootstrap_dofuture.
Resamples subjects, and re-runs the entire subgroup search inside
every replicate, so the replicate may select a different subgroup than the
original analysis did. That is what lets it correct the selection-induced
bias in the reported treatment effect, and supply an
infinitesimal-jackknife interval. It is the reference method, and it is
expensive: cost scales with nb_boots times the cost of one search.
Multiplier resampling (MR) — fs_mr_inference, turned
on by mr_inference = TRUE. Resamples nothing; it perturbs the
per-subject influence contributions (dfbeta) of a single set of
fits with random mean-zero multipliers, and re-applies the selection rule on
each perturbed draw. It corrects the same selection bias FB corrects and
produces the same kind of interval, to leading order, at a small fraction of
the cost — no refits. MR is post-selection inference on a
completed analysis: it runs after the subgroup has been chosen and
cannot change which subgroup is identified. Running with
mr_inference = TRUE and FALSE gives identical
sg.harm, df.est and max_sg_est. MR approximates FB;
it does not replace it.
Harm confirmation. Once MR has produced a de-biased estimate,
mr_inference_args$confirm_rule decides whether that estimate still
indicates harm:
"point"(default)Compares the de-biased point estimate to
t_confirm."ci"Compares the one-sided 95% selection-adjusted lower bound to
t_confirm. Strictly the stronger requirement, so anything it confirms"point"also confirms.
t_confirm is the threshold those rules compare against, on
the effect scale (HR, OR, RR, IRR; or RD, MD for differences) —
not the working log scale used internally. Left NULL it resolves to
the null effect: 1 for ratio measures, 0 for differences. It
is deliberately set near the null rather than at hr.threshold,
because the de-biasing correction over-shrinks genuine effects, so
re-applying the screening threshold to a corrected estimate would discard
true findings.
mr_harm_confirmed carries the answer, and is three-valued:
TRUE = MR ran and harm is confirmed; FALSE = MR ran and harm
is not confirmed; NA = MR did not run, or could not be computed.
NA is not evidence against harm. It arises when
mr_inference = FALSE (the default), when no subgroup was identified,
or when the MR step was skipped or failed for the outcome/consistency
combination in use. Because isTRUE(NA) is FALSE, code that
reads this field with isTRUE() alone silently reports "harm not
confirmed" for an analysis that was never performed; branch on
is.na() first.
References
FDA Guidance for Industry: Enrichment Strategies for Clinical Trials
Athey & Imbens (2016). Recursive partitioning for heterogeneous causal effects. PNAS.
Wager & Athey (2018). Estimation and inference of heterogeneous treatment effects using random forests. JASA.
See Also
subgroup.consistency for consistency evaluation details
forestsearch_bootstrap_dofuture for the full bootstrap (FB)
fs_mr_inference for multiplier resampling (MR), the
mr_inference = TRUE step
mr_estimates_table to render the MR estimates
fs_fdr_report to sweep harm-confirmation thresholds
(c_confirm) and measure what confirmation buys
forestsearch_Kfold for cross-validation
Package website: https://larry-leon.github.io/forestsearch/
Source code: https://github.com/larry-leon/forestsearch
Examples
## Not run:
# Example 1: Standard analysis (backward compatible)
result <- forestsearch(
df.analysis = trial_data,
sg_focus = "hr",
hr.threshold = 1.25,
pconsistency.threshold = 0.90,
fs.splits = 400,
details = TRUE
)
# Example 2: Fast exploratory analysis with two-stage
result_fast <- forestsearch(
df.analysis = trial_data,
sg_focus = "maxSG",
hr.threshold = 1.15,
pconsistency.threshold = 0.85,
fs.splits = 500,
use_twostage = TRUE,
details = TRUE
)
# Example 3: Two-stage with custom parameters
result_custom <- forestsearch(
df.analysis = trial_data,
sg_focus = "hr",
hr.threshold = 1.3,
pconsistency.threshold = 0.95,
fs.splits = 600,
use_twostage = TRUE,
twostage_args = list(
n.splits.screen = 50,
batch.size = 25,
conf.level = 0.99
),
parallel_args = list(plan = "multisession", workers = 4),
details = TRUE
)
# Example 4: Binary outcome with OR thresholds
result_binary <- forestsearch(
df.analysis = trial_data,
outcome_type = "binary",
effect_measure = "OR",
effect.threshold = 1.5,
consistency.threshold = 1.3,
pconsistency.threshold = 0.90
)
# Example 5: Binary outcome with RD thresholds
result_rd <- forestsearch(
df.analysis = trial_data,
outcome_type = "binary",
effect_measure = "RD",
effect.threshold = 0.07, # 7 pct-point harm
consistency.threshold = 0.03 # 3 pct-point per split
)
## End(Not run)
ForestSearch K-Fold Cross-Validation
Description
This function assesses the stability and reproducibility of ForestSearch subgroup identification through cross-validation. For each fold:
Train ForestSearch on (K-1) folds
Apply the identified subgroup to the held-out fold
Compare predictions to the original full-data analysis
Usage
forestsearch_Kfold(
fs.est,
Kfolds = .fs_cv_nobs(fs.est),
seedit = 8316951L,
parallel_args = list(plan = "multisession", workers = 6, show_message = TRUE),
sg0.name = "Not recommend",
sg1.name = "Recommend",
details = FALSE,
mr_in_replicates = FALSE
)
Arguments
fs.est |
List. ForestSearch results object from |
Kfolds |
Integer. Number of folds (default: |
seedit |
Integer. Random seed for fold assignment (default: 8316951). |
parallel_args |
List. Parallelization configuration with elements:
|
sg0.name |
Character. Label for subgroup 0 (default: "Not recommend"). |
sg1.name |
Character. Label for subgroup 1 (default: "Recommend"). |
details |
Logical. Print progress details (default: FALSE). |
mr_in_replicates |
Logical. Whether multiplier resampling (MR) runs
inside each CV fold. Default
The retained objects are not aggregable into an estimate for the original analysis. Each fold re-runs the search on its training split, selects its own subgroup, and de-biases that subgroup – so the collection describes MR's behaviour across fold-specific selections, not a better estimate of the original quantity. Averaging them, or pooling their intervals, is an error. |
Details
Performs K-fold cross-validation for ForestSearch, evaluating subgroup identification and agreement between training and test sets.
Value
List with components:
- resCV
Data frame with CV predictions for each observation
- cv_args
Arguments used for CV ForestSearch calls
- timing_minutes
Execution time in minutes
- prop_SG_found
Percentage of folds where a subgroup was found
- sg_analysis
Original subgroup definition from full-data analysis
- sg0.name, sg1.name
Subgroup labels
- Kfolds
Number of folds used
- sens_summary
Named vector of sensitivity metrics (sens_H, sens_Hc, ppv_H, ppv_Hc)
- find_summary
Named vector of subgroup-finding metrics (Any, Exact, etc.)
- mr_replicates
NULLunder the defaultmr_in_replicates = FALSE. Otherwise a list of lengthKfolds, elementkholding foldk'smr_inferenceobject (orNULL), each computed against that fold's candidate family and selected subgroup. Not aggregable – see themr_in_replicatesparameter.
Cross-Validation Types
-
Leave-One-Out (LOO): When
Kfolds = nrow(df), each observation is held out once. Most thorough but computationally intensive. -
K-Fold: When
Kfolds < nrow(df), data is split into K roughly equal folds. Good balance of bias-variance tradeoff.
Output Metrics
The returned resCV data frame contains:
-
treat.recommend: Prediction from CV model -
treat.recommend.original: Prediction from full-data model -
cvindex: Fold assignment -
sg1,sg2: Subgroup definitions found in each fold
See Also
forestsearch for initial subgroup identification
forestsearch_KfoldOut for summarizing CV results
forestsearch_tenfold for repeated K-fold simulations
Examples
## Not run:
# Run initial ForestSearch
fs_result <- forestsearch(
df.analysis = trial_data,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
confounders.name = c("age", "biomarker")
)
# Run 10-fold cross-validation
cv_results <- forestsearch_Kfold(
fs.est = fs_result,
Kfolds = 10,
parallel_args = list(plan = "multisession", workers = 4),
details = TRUE
)
# Summarize results
cv_summary <- forestsearch_KfoldOut(cv_results, outall = TRUE)
## End(Not run)
ForestSearch K-Fold Cross-Validation Output Summary
Description
Summarizes cross-validation results for ForestSearch, including subgroup agreement and performance metrics.
Usage
forestsearch_KfoldOut(res, details = FALSE, outall = FALSE, digits = 4)
Arguments
res |
List. Result object from ForestSearch cross-validation, must contain
elements: |
details |
Logical. Print details during execution (default: FALSE). |
outall |
Logical. If TRUE, returns all summary tables; if FALSE, returns only metrics (default: FALSE). |
digits |
Integer. Decimal places for construction-time numeric
formatting in the per-fold summary tables built when
|
Value
If outall=FALSE, a list with sens_metrics_original and
find_metrics. If outall=TRUE, a list with summary tables and metrics.
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
fs <- forestsearch(df,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
oob <- forestsearch_Kfold(fs.est = fs, Kfolds = nrow(df))
result <- forestsearch_KfoldOut(oob, outall = TRUE)
## End(Not run)
ForestSearch Bootstrap with doFuture Parallelization
Description
Orchestrates bootstrap analysis for ForestSearch using doFuture parallelization. Implements bias correction methods to adjust for optimism in subgroup selection.
Usage
forestsearch_bootstrap_dofuture(
fs.est,
nb_boots,
seed = 8316951L,
details = FALSE,
show_three = FALSE,
parallel_args = list(),
digits = 4,
mr_in_replicates = FALSE
)
Arguments
fs.est |
List. ForestSearch results object from |
nb_boots |
Integer. Number of bootstrap samples (recommend 500-1000). |
seed |
Integer. Random seed for reproducibility of bootstrap sample
generation. Default |
details |
Logical. If |
show_three |
Logical. If |
parallel_args |
List. Parallelization configuration with elements:
If empty list, inherits settings from original forestsearch call. |
digits |
Integer. Decimal places for construction-time numeric
formatting in |
mr_in_replicates |
Logical. Whether multiplier resampling (MR) runs
inside each bootstrap replicate. Default
The retained objects are not aggregable into an estimate for the
original analysis. Each replicate re-runs the search, selects its own
subgroup, and de-biases that subgroup – so the collection
describes MR's sampling behaviour across resampled selections, not a
better estimate of the original quantity. Averaging them, or pooling
their intervals, is an error. The bias-corrected estimate for the
original analysis is the one this function already returns in
|
Value
An fs_bootstrap object (a list with class
c("fs_bootstrap", "list")) containing:
- results
Data.table with bias-corrected estimates for each bootstrap iteration
- SG_CIs
List of confidence intervals for H and Hc (raw and bias-corrected)
- FSsg_tab
Formatted table of subgroup estimates
- Ystar_mat
Matrix (nb_boots x n) of bootstrap sample indicators
- H_estimates
Detailed estimates for subgroup H
- Hc_estimates
Detailed estimates for subgroup Hc
- mr_replicates
NULLunder the defaultmr_in_replicates = FALSE. Otherwise a list of lengthnb_boots, elementbholding replicateb'smr_inferenceobject (orNULLfor a failed or no-subgroup replicate), each computed against that replicate's candidate family and selected subgroup. See themr_in_replicatesparameter for why these must not be averaged.- summary
(If create_summary=TRUE) Enhanced summary with tables and diagnostics
- nb_boots
Integer. Number of bootstrap iterations requested. Used by
print.fs_bootstrapto compute identification percentages without re-inspectingresults. Added in v0.2.0.- original_sg
Character vector. The subgroup identified by the primary
forestsearchcall (fs.est$sg.harm), carried forward so thatprint.fs_bootstrapcan report exact- and partial-match rates without retaining a reference tofs.est. Added in v0.2.0.- outcome_type
Character. Outcome type of the underlying analysis: one of
"survival","binary","continuous", or"count". Added in v0.2.0.- effect_measure
Character. Effect measure label (
"HR"for survival;"OR","RR","RD","IRR","IRD", or"MD"for GLM outcomes). Added in v0.2.0.- est.scale
Character. Estimation scale for confidence intervals (
"hr"or"1/hr"). Added in v0.2.0.
Bias Correction Methods
Two bias correction approaches are implemented:
-
Method 1 (Simple Optimism):
H_{adj1} = H_{obs} - (H^*_{*} - H^*_{obs})where
H^*_{*}is the new subgroup HR on bootstrap data andH^*_{obs}is the new subgroup HR on original data. -
Method 2 (Double Bootstrap):
H_{adj2} = 2 \times H_{obs} - (H_{*} + H^*_{*} - H^*_{obs})where
H_{*}is the original subgroup HR on bootstrap data.
Variable Naming Convention
-
H: Original subgroup (harm/questionable, treat.recommend == 0) -
Hc: Complement subgroup (recommend, treat.recommend == 1) -
_obs: Estimate from original data -
_star: Estimate from bootstrap data -
_biasadj_1: Bias correction method 1 -
_biasadj_2: Bias correction method 2
Valid for every identifier configuration
Unlike multiplier resampling, the full bootstrap places no alignment requirement on the identifier. It replays the entire pipeline on each resample – refitting the front ends, re-enumerating the candidate family and re-running selection – rather than linearizing the selection event, so neither of MR's two conditions (selection ranking on the inferential coefficient, and a fixed candidate family) bears on its validity. In the equivalence result those conditions state, the bootstrap is the reference standard, not a party to it.
It is therefore valid for all six identifier configurations, including the
three that rank on a native statistic: DINA under
dina_select_statistic = "dina", GRF under
grf_select_statistic = "dr", and GRF under
grf_selection = "tree". For those, it correctly estimates the
selection bias that ranking on the native statistic induces on the
reported effect \hat\beta. That is a coherent quantity, and the
estimate of it is valid – it simply sits outside the framework the
manuscript develops, which is about identifiers whose selection map and
reported effect are the same functional.
The one place an alignment check does apply here is
mr_in_replicates = TRUE, which asks for MR inside each
replicate; see that parameter.
Performance
Typical runtime: 1-5 seconds per bootstrap iteration. For 1000 bootstraps with 6 workers, expect 3-10 minutes total. Memory usage scales with dataset size and number of workers.
Requirements
Original
fs.estmust have identified a valid subgroupRequires packages:
data.table,foreach,doFuture,survivalFor plots: requires
ggplot2
See Also
forestsearch for initial subgroup identification, and its
vocabulary section for how the full bootstrap (FB) implemented here
relates to multiplier resampling (MR)
fs_mr_inference for MR, the fast refit-free approximation to
this function
mr_estimates_table to render MR estimates in this function's
table layout
fs_fdr_report for harm-confirmation threshold sweeps
bootstrap_results for the core bootstrap worker function
build_cox_formula for Cox formula construction
fit_cox_models for Cox model fitting
Examples
## Not run:
# Run ForestSearch
fs_result <- forestsearch(
df.analysis = mydata,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
confounders.name = c("age", "sex", "stage")
)
# Run bootstrap with bias correction
boot_results <- forestsearch_bootstrap_dofuture(
fs.est = fs_result,
nb_boots = 1000,
parallel_args = list(
plan = "multisession",
workers = 6,
show_message = TRUE
),
create_summary = TRUE,
create_plots = TRUE
)
# View results
print(boot_results$FSsg_tab)
print(boot_results$summary$table)
# Check success rate
mean(!is.na(boot_results$results$H_biasadj_2))
## End(Not run)
ForestSearch Repeated K-Fold Cross-Validation
Description
This function performs multiple independent K-fold cross-validations to assess the variability in subgroup identification. Each simulation:
Randomly shuffles the data
Performs K-fold CV
Records sensitivity and agreement metrics
Results are summarized across all simulations.
Usage
forestsearch_tenfold(
fs.est,
sims,
Kfolds = 10,
details = TRUE,
seed = 8316951L,
parallel_args = list(plan = "multisession", workers = 6, show_message = TRUE),
keep_resCV = FALSE,
mr_in_replicates = FALSE
)
Arguments
fs.est |
List. ForestSearch results object from |
sims |
Integer. Number of simulation repetitions. |
Kfolds |
Integer. Number of folds per simulation (default: 10). |
details |
Logical. Print progress details (default: TRUE). |
seed |
Integer. Base random seed for fold shuffling. Default 8316951L. Each simulation uses seed + 1000 * ksim for reproducibility. |
parallel_args |
List. Parallelization configuration (elements
|
keep_resCV |
Logical. If |
mr_in_replicates |
Logical. Whether multiplier resampling (MR) runs
inside each CV fold. Default
The retained objects are not aggregable into an estimate for the original analysis. Each fold re-runs the search on its training split, selects its own subgroup, and de-biases that subgroup – so the collection describes MR's behaviour across fold-specific selections, not a better estimate of the original quantity. Averaging them, or pooling their intervals, is an error. |
Details
Runs repeated K-fold cross-validation simulations for ForestSearch and summarizes subgroup identification stability across repetitions.
Value
List with components:
- sens_summary
Named vector of median sensitivity metrics across simulations
- find_summary
Named vector of median subgroup-finding metrics
- sens_out
Matrix of sensitivity metrics (sims x metrics)
- find_out
Matrix of finding metrics (sims x metrics)
- fold_summary
Data frame with one row per (sim, fold) combination. Columns:
sim,fold,n_test,sg1,sg2,grf_cuts,pconsistency,training_fs_hr,n_candidates_evaluated,any_found. Thegrf_cutscolumn records the GRF policy-tree cut expressions returned for each training fold (collapsed with " | " when multiple cuts are returned;NAwhen GRF was not used, failed, or returned no cuts). Thepconsistencycolumn records the consistency probability (Pcons) achieved by the identified subgroup on each fold (NA_real_when no subgroup was identified or the training ForestSearch call errored). Thetraining_fs_hrcolumn records the in-sample effect estimate for the identified subgroup on the training fold, on its NATURAL SCALE regardless of outcome type (hazard ratio for survival; odds ratio, rate ratio, or risk ratio for GLM ratio measures after internal exponentiation; mean/risk difference for GLM additive measures); optimistically biased relative to any independent estimate and surfaced for diagnostic comparison only, withNA_real_when no subgroup was identified. Then_candidates_evaluatedcolumn records the integer count of candidate subgroups actually evaluated for consistency on this training fold (populated whenever the consistency stage ran, even if zero candidates met the threshold and thus no subgroup was identified;NA_integer_when the training call errored or consistency evaluation did not run). Always returned (compact; cheap). Lets you tabulate, for example, which subgroup was identified in each fold of each simulation, the empirical distribution of GRF cut choices across the full sim x fold grid, the relationship between GRF's cut and the final identified subgroup, or the near-miss consistency values among folds that did not surface a subgroup.- resCV_all
List of length
sims; element i is the per-subjectresCVdata frame from simulation i. Only populated whenkeep_resCV = TRUE; otherwiseNULL.- timing_minutes
Total execution time
- sims
Number of simulations run
- Kfolds
Number of folds per simulation
- mr_replicates
NULLunder the defaultmr_in_replicates = FALSE. Otherwise a nested list: one element per successful simulation (namedsim<k>), each a list of lengthKfoldsholding that fold'smr_inferenceobject (orNULL), computed against that fold's candidate family and selected subgroup. Not aggregable – see themr_in_replicatesparameter.
Parallelization Strategy
Unlike the single K-fold function which parallelizes across folds, this function parallelizes across simulations for better efficiency when running many repetitions. Each simulation runs its K-fold CV sequentially.
See Also
forestsearch_Kfold for single K-fold CV
forestsearch_KfoldOut for summarizing CV results
Examples
## Not run:
# Run 100 repetitions of 10-fold CV
tenfold_results <- forestsearch_tenfold(
fs.est = fs_result,
sims = 100,
Kfolds = 10,
parallel_args = list(plan = "multisession", workers = 6),
details = TRUE
)
# View summary
print(tenfold_results$sens_summary)
print(tenfold_results$find_summary)
## End(Not run)
Format Confidence Interval for Estimates
Description
Formats confidence interval for estimates.
Usage
format_CI(estimates, col_names, digits = 4)
Arguments
estimates |
Data frame or data.table of estimates. |
col_names |
Character vector of column names for estimate, lower, upper. |
digits |
Integer. Decimal places for numeric formatting. Default 4
to provide round-down headroom for downstream display reformatting
(e.g., |
Value
Character string formatted as \"estimate (lower, upper)\".
Examples
library(data.table)
est <- data.table(hr = 1.58, lower = 0.86, upper = 2.90)
format_CI(est, col_names = c("hr", "lower", "upper"))
Format Bootstrap Diagnostics Table with gt
Description
Creates a publication-ready diagnostics table from bootstrap results.
Usage
format_bootstrap_diagnostics_table(
diagnostics,
nb_boots,
results,
H_estimates = NULL,
Hc_estimates = NULL,
effect_label = "HR"
)
Arguments
diagnostics |
List. Diagnostics information from summarize_bootstrap_results() |
nb_boots |
Integer. Number of bootstrap iterations |
results |
Data.table. Bootstrap results with bias-corrected estimates |
H_estimates |
List. H subgroup estimates |
Hc_estimates |
List. Hc subgroup estimates |
Value
A gt table object
Format Bootstrap Results Table with gt
Description
Creates a publication-ready table from ForestSearch bootstrap results, with bias-corrected confidence intervals, informative formatting, and optional subgroup definition footnote.
Usage
format_bootstrap_table(
FSsg_tab,
nb_boots,
est.scale = "hr",
boot_success_rate = NULL,
sg_definition = NULL,
title = NULL,
subtitle = NULL
)
Arguments
FSsg_tab |
Data frame or matrix from forestsearch_bootstrap_dofuture()$FSsg_tab |
nb_boots |
Integer. Number of bootstrap iterations performed |
est.scale |
Character. "hr" or "1/hr" for effect scale |
boot_success_rate |
Numeric. Proportion of bootstraps that found subgroups |
sg_definition |
Character. Subgroup definition string to display as footnote (e.g., "{age>=50} & {nodes>=3}"). If NULL, no subgroup footnote is added. |
title |
Character. Custom title (optional) |
subtitle |
Character. Custom subtitle (optional) |
Value
A gt table object
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
fs_bc <- forestsearch_bootstrap_dofuture(fs, nb_boots = 100)
summaries <- summarize_bootstrap_results(fs$sg.harm, fs_bc)
format_bootstrap_table(summaries$table_data)
## End(Not run)
Format Bootstrap Timing Table with gt
Description
Creates a publication-ready timing summary table from bootstrap results.
Usage
format_bootstrap_timing_table(timing_list, nb_boots, boot_success_rate)
Arguments
timing_list |
List. Timing information from summarize_bootstrap_results()$timing |
nb_boots |
Integer. Number of bootstrap iterations |
boot_success_rate |
Numeric. Proportion of successful bootstraps |
Value
A gt table object
Format Continuous Variable Definition for Display
Description
Format Continuous Variable Definition for Display
Usage
format_continuous_definition(var_data, cut_spec, var_name)
Format ForestSearch Details Output for Beamer Two-Column Display
Description
Captures forestsearch(details = TRUE) console output and splits it
into two columns for readable beamer slides. Left column shows variable
selection (GRF, LASSO, candidate factors); right column shows subgroup
search, consistency evaluation, and results.
Usage
format_fs_details(
fs_output,
split_after = "Candidate factors",
fontsize = "scriptsize",
col_widths = c(0.48, 0.52),
max_width = 48
)
Arguments
fs_output |
Character vector of captured output lines from
|
split_after |
Character string (regex). The output is split after the
block matching this pattern. Default: |
fontsize |
Character. LaTeX font size for the output text.
One of |
col_widths |
Numeric vector of length 2. Column widths as fractions
of |
max_width |
Integer. Maximum character width per line before wrapping. Long lines are wrapped at comma or space boundaries with a 4-space continuation indent. Default: 48 (suitable for half-slide columns at scriptsize). |
Value
Invisibly returns a list with left and right
character vectors. Side effect: emits LaTeX via cat() for use
in a chunk with results='asis'.
Quarto Setup
No special LaTeX packages required. Works in any beamer frame
without the fragile option.
Usage
In a Quarto beamer chunk with results='asis' and
echo=FALSE, first capture the forestsearch output with
capture.output(), then call format_fs_details(fs_output).
Examples
## Not run:
fs_output <- capture.output(
fs_res <- forestsearch(
df = dat, outcome.name = "time", event.name = "status",
treat.name = "trt", id.name = "id",
confounders.name = confs, details = TRUE
)
)
# Default: split after candidate factors
format_fs_details(fs_output, fontsize = "tiny")
# Custom split: after filtering summary
format_fs_details(fs_output, split_after = "Found.*candidate")
## End(Not run)
Format Operating Characteristics Results as GT Table
Description
Creates a formatted gt table from simulation operating characteristics results.
Usage
format_oc_results(
results,
analyses = NULL,
metrics = "all",
digits = 3,
digits_hr = 3,
title = "Operating Characteristics Summary",
subtitle = NULL,
use_gt = TRUE,
effect_measure = NULL,
dgm = NULL,
subgroup_notation = c("harm", "benefit")
)
Arguments
results |
data.table or data.frame. Simulation results from
|
analyses |
Character vector. Analysis methods to include. Default: NULL (all analyses in results) |
metrics |
Character vector. Metrics to display. Options include: "detection", "classification", "hr_estimates", "ahr_estimates", "cde_estimates", "subgroup_size", "all". Default: "all" |
digits |
Integer. Decimal places for proportions. Default: 3 |
digits_hr |
Integer. Decimal places for hazard ratios. Default: 3 |
title |
Character. Table title. Default: "Operating Characteristics Summary" |
subtitle |
Character or |
use_gt |
Logical. Return gt table if TRUE, data.frame if FALSE. Default: TRUE |
effect_measure |
Character or |
dgm |
Optional DGM object ( |
subgroup_notation |
Character. |
Details
The function summarizes simulation results across multiple metrics:
-
Found: Proportion of simulations finding a subgroup (any.H)
-
Classification: Sen, spec, PPV, NPV
-
HR Estimates: Mean Cox hazard ratios in true (H) and identified (H-hat) subgroups and their complements
-
AHR Estimates: Mean average hazard ratios (from loghr_po) in true and identified subgroups
-
CDE Estimates: Controlled direct effects (from theta_0/theta_1) in true and identified subgroups
-
Subgroup Size: Average, min, max sizes
Column notation aligns with build_estimation_table and
Leon et al. (2024): H = true (oracle) subgroup, H-hat =
identified subgroup. The asterisk (*) is reserved for bootstrap
bias-corrected estimates and is not used in this summary table.
Value
A gt table object (if use_gt = TRUE and gt package available) or data.frame
Examples
## Not run:
# format_oc_results() is called by summarize_simulation_results().
# See run_simulation_analysis() for the standard entry point.
## End(Not run)
Format results for subgroup summary
Description
Formats results for subgroup summary table.
Usage
format_results(
subgroup_name,
n,
n_treat,
d,
m1,
m0,
drmst,
hr,
hr_a = NA,
hr_po = NA,
return_medians = TRUE
)
Arguments
subgroup_name |
Character. Subgroup name. |
n |
Character. Sample size. |
n_treat |
Character. Treated count. |
d |
Character. Event count. |
m1 |
Numeric. Median or RMST for treatment. |
m0 |
Numeric. Median or RMST for control. |
drmst |
Numeric. RMST difference. |
hr |
Character. Hazard ratio (formatted). |
hr_a |
Character. Adjusted hazard ratio (optional). |
hr_po |
Numeric. Potential outcome hazard ratio (optional). |
return_medians |
Logical. Use medians or RMST. |
Value
Character vector of results.
Format Search Results
Description
Format Search Results
Usage
format_search_results(
results_list,
Z,
details,
t.sofar,
L,
max_count,
filter_counts = NULL
)
Arguments
results_list |
List of result rows |
Z |
Matrix of factor indicators |
details |
Logical. Print details |
t.sofar |
Numeric. Time elapsed |
L |
Integer. Number of factors |
max_count |
Integer. Maximum combinations |
filter_counts |
List. Counts at each filtering stage (optional) |
Format Subgroup Summary Tables with gt
Description
Creates publication-ready gt tables for bootstrap subgroup analysis
Usage
format_subgroup_summary_tables(subgroup_summary, nb_boots)
Arguments
subgroup_summary |
List from summarize_bootstrap_subgroups() |
nb_boots |
Integer. Number of bootstrap iterations |
Value
List of gt table objects
Approximate Procedure-Level False Positive Rate
Description
Computes an analytical approximation to the ForestSearch false positive
rate under the null hypothesis of homogeneous treatment effect. The
calculation accounts for (i) the sampling variability of the effect
estimate in a subgroup of size n_min, (ii) the number of
candidate subgroups searched (driven by the covariate space and
maxk), and (iii) the screening threshold c1.
Usage
fpr_approx(
effect_true = 0.65,
p_event = 0.4,
n_min = 60L,
c1 = 1.25,
n_confounders_cont = 6L,
n_confounders_bin = 6L,
cont_cutoff = 4L,
maxk = 2L,
effect_type = c("ratio", "difference"),
treat_alloc = 0.5,
rho = 0.3,
n_eff = NULL,
thresholds = NULL,
verbose = TRUE
)
Arguments
effect_true |
Numeric. True effect under H0 on the ratio scale (e.g., OR or HR). Default 0.65 (protective treatment). |
p_event |
Numeric in (0, 1). Overall event rate (control arm baseline rate for binary; used to compute SE). Default 0.40. |
n_min |
Integer. Minimum subgroup size ( |
c1 |
Numeric. The screening threshold on a candidate
subgroup's own effect – |
n_confounders_cont |
Integer. Number of continuous confounders. Default 6. |
n_confounders_bin |
Integer. Number of binary confounders. Default 6. |
cont_cutoff |
Integer. Number of cut points per continuous
covariate ( |
maxk |
Integer. Maximum interaction depth. 1 = single-variable subgroups only; 2 = up to two-variable interactions. Default 2. |
effect_type |
Character. |
treat_alloc |
Numeric in (0, 1). Treatment allocation fraction. Default 0.5 (1:1 randomisation). |
rho |
Numeric in [0, 1). Assumed average pairwise correlation
between candidate test statistics, used for the correlation-adjusted
FPR estimate. Default 0.30. Set to 0 for the independence
(upper-bound) estimate only. Ignored when |
n_eff |
Numeric or |
thresholds |
Numeric vector. Additional thresholds to evaluate
besides |
verbose |
Logical. Print a formatted summary. Default
|
Details
The approximation assumes:
Normal approximation to the log-OR (or log-HR) distribution.
Independent candidates (upper bound; true FPR is lower due to overlap). A correlation-adjusted estimate is also returned.
Equal allocation to treatment arms within each subgroup.
No consistency check adjustment (the reported FPR is the screening-stage FPR before
pconsistency.thresholdfiltering).
Value
A list (class "fpr_approximation") with components:
se_log_effectSE of log(effect) in a subgroup of size
n_min.n_candidatesTotal number of candidate subgroups searched.
n_singleNumber of single-variable candidates.
n_pairsNumber of two-variable candidates (0 if
maxk == 1).per_candidateNamed numeric vector: P(effect > thr) for each threshold evaluated.
fpr_indepNamed numeric vector: procedure-level FPR assuming independent candidates, for each threshold.
fpr_adjustedNamed numeric vector: correlation-adjusted FPR, for each threshold.
paramsList of input parameters.
Examples
# ACTG175 binary harm search
fpr_approx(effect_true = 0.65, p_event = 0.40, n_min = 60, c1 = 1.25)
# Tighter threshold
fpr_approx(effect_true = 0.65, p_event = 0.40, n_min = 60, c1 = 2.0)
# Larger minimum subgroup
fpr_approx(effect_true = 0.65, p_event = 0.40, n_min = 100, c1 = 1.25)
Direct Simulation-Based FPR Calibration at a Specific Threshold Pair
Description
Estimates the procedure-level false-positive rate (FPR) for a
forestsearch analysis at a specific (c_1, c_2)
threshold pair by running nsims null simulations. No power-law
model or threshold-independence assumption is required.
Usage
fpr_calibration(
c1,
c2,
df.analysis = NULL,
fs_params = list(),
nsims = 200L,
nsims_verify = 0L,
null_fn = NULL,
n_sample = NULL,
n_min = 60L,
outcome_type_approx = NULL,
d_eff_approx = NULL,
prop_cens = 0,
p_event = 0.5,
sigma_y = 1,
sim_seed = 42L,
n_workers = 1L,
quiet = FALSE
)
Arguments
c1 |
Numeric scalar. The screening threshold on a candidate
subgroup's own effect. |
c2 |
Numeric scalar. The per-split effect threshold. |
df.analysis |
Data frame. The analyst's dataset. Used directly in
permutation mode; also used to extract N for the per-subgroup
approximation. Can be |
fs_params |
Named list. All forestsearch() arguments except
|
nsims |
Integer. Number of null replicates. Default 200.
At FPR ~ 10%, |
nsims_verify |
Integer. Number of independent verification
replicates. Uses a separate seed range from the calibration batch
to provide an honest check of the corrected FPR. Default 0 (skip
verification). Set to |
null_fn |
Function or |
n_sample |
Integer or |
n_min |
Integer. Minimum subgroup size for the per-subgroup
approximation |
outcome_type_approx |
Character. One of |
d_eff_approx |
Numeric or |
prop_cens |
Numeric. Censoring proportion for survival outcomes.
Used only when |
p_event |
Numeric. Event probability for binary outcomes.
Used only when |
sigma_y |
Numeric. Residual SD for continuous outcomes. Default 1. |
sim_seed |
Integer. Base random seed. Each replicate uses
|
n_workers |
Integer. Parallel workers for the foreach loop.
Default 1 (sequential). Set to |
quiet |
Logical. Suppress progress messages. Default |
Details
Two null-generation modes are supported:
null_fnprovidedEach replicate draws a fresh null dataset from
null_fn(seed). Use this when you have a parametric data generating mechanism (DGM) or want to mimic a specific study design.df.analysisonly (permutation mode)Each replicate permutes the treatment column within the analyst's actual dataset. This preserves the observed covariate and outcome distributions and is the natural H0 reference for a real analysis. Requires the
treat.namecolumn to be present indf.analysis.
Permutation mode
Under the null hypothesis of no subgroup heterogeneity, the treatment
assignment is exchangeable with the observed outcomes. Permuting
treat.name within each replicate breaks any true treatment-by-
subgroup interaction while preserving the marginal distributions of all
other variables. This is a model-free H0 reference that adapts
automatically to the actual dataset structure.
P1 approximation
The asymptotic per-subgroup P_1 at (c1, c2) is computed
via compute_detection_probability_glm() using the effective
information d_{\text{eff}} at n.min. The ratio
L_{\text{eff}} = \log(1 - \hat{\text{FPR}}) / \log(1 - P_1)
is the implied multiplicity factor specific to this threshold pair and
dataset — it replaces the universal power-law model and does not assume
threshold-independence.
Value
An object of class "fpr_calibration" with components:
fprNumeric. Estimated procedure-level FPR.
fpr_ciNumeric vector of length 2. 95% Wilson CI.
n_foundInteger. Number of replicates where a subgroup was found.
nsimsInteger. Total replicates attempted.
n_failedInteger. Replicates that errored.
P1Numeric. Per-subgroup approximation
P_1at(c1, c2).d_effNumeric. Effective information used for
P_1.L_eff_impliedNumeric. Implied
L_{\text{eff}}from the simulation result:\log(1 - \text{FPR}) / \log(1 - P_1).c1, c2Numeric. The threshold pair.
NInteger. Sample size.
modeCharacter.
"permutation"or"dgm".call_argsList. Key call arguments for reproducibility.
See Also
forestsearch,
compute_detection_probability_glm,
d_eff_binary, d_eff_survival
Examples
## Not run:
# ---- Binary example (permutation mode) ----
set.seed(1)
df <- data.frame(
id = 1:500, treat = rbinom(500, 1, 0.5),
bm1 = as.factor(rbinom(500, 1, 0.7)),
bm2 = as.factor(rbinom(500, 1, 0.5)),
age = round(rnorm(500, 55, 10)),
y = rbinom(500, 1, 0.30) # null: no treatment effect
)
fs_params <- list(
confounders.name = c("bm1", "bm2", "age"),
outcome.name = "y", event.name = "y",
treat.name = "treat", id.name = "id",
outcome_type = "binary", effect_measure = "OR",
adverse_outcome = TRUE,
pconsistency.threshold = 0.90, fs.splits = 400,
n.min = 60, maxk = 2,
use_lasso = TRUE, use_grf = TRUE,
is.RCT = TRUE, details = FALSE, quiet = TRUE,
parallel_args = list(plan = "sequential", workers = 1)
)
res <- fpr_calibration(
c1 = 1.70, c2 = 1.70,
df.analysis = df,
fs_params = fs_params,
nsims = 200,
n_min = 60, outcome_type_approx = "binary", p_event = 0.30,
sim_seed = 42, n_workers = 4
)
print(res)
## End(Not run)
Transfer Corrected FPR Prediction to a Nearby Threshold Pair
Description
Uses the implied L_{\text{eff}} from a calibrated
fpr_calibration object to predict the procedure-level FPR at
(c_1 + \delta_1, c_2 + \delta_2), then optionally runs a direct
simulation at the new pair to validate the prediction.
Usage
fpr_calibration_transfer(
cal_result,
delta = c(0.05, 0.1, 0.2),
delta_c1 = NULL,
delta_c2 = NULL,
nsims_check = 200L,
...
)
Arguments
cal_result |
An object of class |
delta |
Numeric scalar. Added to both |
delta_c1 |
Numeric or |
delta_c2 |
Numeric or |
nsims_check |
Integer. Simulations to run at the new pair for empirical validation. Default 200. Set to 0 to skip. |
... |
Additional arguments passed to |
Value
A data frame with one row per delta value containing:
c1_new, c2_new, P1_new, fpr_pred
(predicted corrected FPR using transferred L_eff_implied),
fpr_sim (direct simulation, if nsims_check > 0),
fpr_sim_lo, fpr_sim_hi (95\
(whether fpr_pred falls within the simulation CI).
Examples
## Not run:
# Calibrate at base pair
cal <- fpr_calibration(c1 = 1.70, c2 = 1.70, null_fn = my_dgm,
n_sample = 500, fs_params = fs_binary,
nsims = 200, n_min = 60,
outcome_type_approx = "binary", p_event = 0.30)
# Predict at base + 0.05, +0.10, +0.20 and validate
transfer_results <- fpr_calibration_transfer(
cal, delta = c(0.05, 0.10, 0.20),
nsims_check = 200, n_workers = 4)
print(transfer_results)
## End(Not run)
Per-Subject Membership Matrix for Pareto Frontier Members
Description
For each candidate subgroup on
fs$grp.consistency$out_sg$pareto_frontier, returns a binary
indicator (one row per subject in fs$df.est, one column per
frontier member) marking which subjects belong to that subgroup.
Usage
frontier_member_flags(fs)
Arguments
fs |
A |
Details
This lets the user compute downstream quantities the frontier table does not surface directly: per-frontier-member sample sizes by covariate stratum, set comparisons between frontier members ("which subjects are in the selected subgroup but not on the runner-up?"), and survival summaries by frontier member.
Value
A list with three elements:
flagsInteger matrix of dimension
n_subjectsxn_frontier_members, with 1 = subject in subgroup, 0 = not.member_labelsCharacter vector of length
n_frontier_members, each entry the "cut1 & cut2 & ..." definition of that frontier member.is_selectedLogical vector of length
n_frontier_members, marking which member is the selected subgroup.
See Also
pareto_frontier_table,
compute_frontier_cis, plot_pareto_frontier.
Examples
## Not run:
fmf <- frontier_member_flags(fs)
# How many subjects in each frontier member?
colSums(fmf$flags)
# Which subjects are unique to the selected subgroup?
sel <- which(fmf$is_selected)
rowSums(fmf$flags) == 1L & fmf$flags[, sel] == 1L
## End(Not run)
Assert the exported selected-subgroup membership matches forestsearch's own
Description
Integrity check for fs_export_maxeff_family(): the materialized membership
column of the selected candidate must equal, row for row, the per-subject
membership forestsearch itself computed
(fit$grp.consistency$sg.harm.id). A mismatch means the label-to-dummy
reconstruction is wrong, so this stops with an error rather than letting a
silently different subgroup be de-biased. fs_to_guohe() runs this check by
default (verify = TRUE).
Usage
fs_assert_membership(exp, fit)
Arguments
exp |
The list returned by |
fit |
The |
Value
TRUE, invisibly, if the memberships agree; otherwise an error is
thrown.
See Also
fs_export_maxeff_family(), fs_to_guohe()
Examples
set.seed(11)
n <- 200
df <- data.frame(id = seq_len(n), treat = rbinom(n, 1, 0.5),
age = round(rnorm(n, 55, 12)),
bm = factor(rbinom(n, 1, 0.5)))
harm <- df$age <= 50 & df$bm == "1"
tt <- rexp(n, 0.05 * exp(log(2.5) * df$treat * harm))
df$time <- pmin(tt, 60)
df$event <- as.integer(tt <= 60)
fit <- forestsearch(
df.analysis = df,
outcome.name = "time", event.name = "event",
treat.name = "treat", id.name = "id",
confounders.name = c("age", "bm"),
sg_focus = "maxeff", use_grf = FALSE, use_lasso = FALSE,
fs.splits = 50, n.min = 30, d0.min = 5, d1.min = 5,
hr.threshold = 1.1, hr.consistency = 1.0,
pconsistency.threshold = 0.5, maxk = 2,
details = FALSE, plot.sg = FALSE,
parallel_args = list(plan = "sequential", workers = 1L,
show_message = FALSE)
)
exp_fam <- fs_export_maxeff_family(fit)
fs_assert_membership(exp_fam, fit)
Attach beta(Hhat) targets to a results frame
Description
The one call each engine adds to its run_cell(), immediately before the
bundle is assembled. Runs in the main process, after the replicate loop, so
the evaluation frame stays in one place.
Usage
fs_attach_betaHhat(
results,
frame,
focus,
outcome_type = c("survival", "binary", "continuous", "count"),
effect_measure = NULL,
outcome.name = "y_sim",
event.name = "event_sim",
treat.name = "treat_sim",
sg_def_structs = NULL
)
Arguments
results |
Data frame carrying an |
frame |
Data frame. The fixed evaluation population. |
focus, outcome_type, effect_measure, outcome.name, event.name, treat.name |
Passed to |
sg_def_structs |
Optional named list of structured GRF definitions. |
Value
results with betaHhat_H, betaHhat_Hc, betaHhat_status and
nH_eval / nHc_eval columns added, and counters attached.
Replicates with no realized rule
An undetected replicate is not a missing measurement – it is a run in which
the whole population is the complement. It therefore receives the
no-subgroup record fs_betaHhat_one() already returns: nH_eval = 0,
nHc_eval = nrow(frame), betaHhat_Hc the ITT effect, betaHhat_H NA
(an empty region has no target), and status = "ok".
This is the partition invariant reaching this layer: nH_eval + nHc_eval
now equals nrow(frame) on every row that is not "unresolved", not
only on rows carrying a rule. Before this change such rows were all-NA,
which silently excluded them from any consumer keying on a finite target.
Their count is still reported separately as n_reps_undetected.
Resolution counters from a beta(Hhat) table or attached results
Description
Resolution counters from a beta(Hhat) table or attached results
Usage
fs_betaHhat_counts(x)
Arguments
x |
Output of |
Value
A named list of counters, or NULL if none are attached.
Refuse to print coverage targets computed on different denominators
Description
.coverage_meta()-style assembly computes
ok <- is.finite(target) & is.finite(lo) & is.finite(hi) and
n_eff <- sum(ok), so a non-finite target is dropped with no error and no
warning. Coverage is then computed over the replicates whose targets
happened to score, which is not a random subset: whatever made a target
non-finite is a property of the realized rule.
Usage
fs_betaHhat_neff_parity(
cov_df,
strict = TRUE,
target_beta = "C_betaHhat",
target_ref = "C_dagger"
)
Arguments
cov_df |
Coverage table with |
strict |
Logical. |
target_beta, target_ref |
Character. The target to check and the reference to check it against. |
Details
Because a reference target that is a finite scalar is dropped only when the
interval itself is non-finite, n_eff parity between the two is the exact
check for a target-specific loss, and it is free – n_eff is already in
the table.
Value
cov_df, invisibly.
Two causes, deliberately not distinguished
A parity break can mean an unresolvable rule, or a genuinely degenerate
region whose per-family guard returned NA by design. Both make the two
coverage figures incomparable as printed, so both stop. Use
fs_betaHhat_neff_report() on the bundle to tell them apart before
deciding.
Which rules produced NA beta(Hhat) targets, and why
Description
Diagnostic companion to fs_betaHhat_neff_parity(). Separates the two
causes of a parity break: a rule the evaluation frame cannot express, and a
region small enough for a per-family guard to fire.
Usage
fs_betaHhat_neff_report(bundle, block = c("H", "Hc"))
Arguments
bundle |
A results data frame, a bundle list with a |
block |
Character. |
Value
A data frame of distinct offending rules with is_disjunction,
nH_eval and status where available, or NULL when there are none.
Conditional estimand beta(Hhat) for one realized rule
Description
Computes the population value of the standard within-subgroup analysis on a realized selection region and its complement, on a fixed evaluation frame. This is the conditional estimand of León and Anderson (2026, Theorem 2), evaluated at the rule a detector actually returned – not the marginal true-subgroup target.
Usage
fs_betaHhat_one(
rule,
frame,
focus,
outcome_type = c("survival", "binary", "continuous", "count"),
effect_measure = NULL,
outcome.name = "y_sim",
event.name = "event_sim",
treat.name = "treat_sim",
sg_def_struct = NULL
)
Arguments
rule |
Character. The realized rule: a named |
frame |
Data frame. The fixed evaluation population. |
focus |
Character, required, no default. |
outcome_type |
Character. One of |
effect_measure |
Character. Passed to |
outcome.name, event.name, treat.name |
Character. Column names in
|
sg_def_struct |
Optional structured GRF subgroup definition. When
supplied it takes precedence over |
Details
Membership is resolved once, by the internal .fs_resolve_membership(), and the same
resolution serves every outcome family: membership is a property of the rule
and the frame, not of the effect measure.
Value
A one-row data frame with columns betaHhat_H, betaHhat_Hc,
nH_eval, nHc_eval, status and missing_cols. The column names are
scale-agnostic and identical across families so downstream
paste0("betaHhat_", suffix) scoring is shared verbatim.
The partition invariant
nH_eval + nHc_eval must equal nrow(frame) on every call where the rule
resolves. Hhat and its complement partition the frame by construction, so a
sum below nrow(frame) means some rows fell into neither side – which
happens when a membership vector carries NA. That is a hard error, not a
warning: if membership is incoherent, every number downstream of it is
meaningless.
Unresolvable rules
A rule naming a column the frame lacks has no membership on that frame. The
record is all-NA with status = "unresolved" and missing_cols naming
the columns, rather than a partially-resolved region. Silently dropping the
offending clause returns a finite value for the wrong region.
Deduplicated beta(Hhat) targets over distinct realized rules
Description
Scores each distinct realized rule once. Deduplication is by rule, not by replicate: two replicates landing on the same rule score the same target, and the target does not depend on the trial.
Usage
fs_betaHhat_table(
sg_defs,
frame,
focus,
outcome_type = c("survival", "binary", "continuous", "count"),
effect_measure = NULL,
outcome.name = "y_sim",
event.name = "event_sim",
treat.name = "treat_sim",
sg_def_structs = NULL
)
Arguments
sg_defs |
Character vector of realized rules, one entry per replicate.
|
frame |
Data frame. The fixed evaluation population. |
focus, outcome_type, effect_measure, outcome.name, event.name, treat.name |
Passed to |
sg_def_structs |
Optional named list of structured GRF definitions, keyed by the rule string, used in preference to string parsing. |
Value
A data frame keyed by sg_def, one row per distinct rule, with the
fs_betaHhat_one() schema. Resolution counters are attached as
attributes and retrievable with fs_betaHhat_counts().
Resolution accounting
Six counters are attached to the result and are the fix for targets being
dropped silently. n_eff reported beside a target is what makes a reduced
denominator visible; a coverage figure computed on an unknown fraction of
replicates is not interpretable, and the fraction has to travel with it.
A seventh, n_reps_undetected, counts replicates with no rule at all, so
that n_reps_resolved + n_reps_unresolved + n_reps_undetected closes
against n_reps_total.
theta-dagger at the true subgroup flag, on the scoring frame
Description
The marginal target at the DGM's own harm flag, computed on the same frame
beta(Hhat) is scored on. Use it as a sanity gate: it should reproduce the
DGM's own subgroup effects.
Usage
fs_betaHhat_theta_dagger_check(
frame,
outcome_type = c("survival", "binary", "continuous", "count"),
harm.name = "flag_harm",
outcome.name = "y_sim",
event.name = "event_sim",
treat.name = "treat_sim",
effect_measure = NULL
)
Arguments
frame |
Data frame. The evaluation population, normally from
|
outcome_type |
Character. One of |
harm.name |
Character. The true-subgroup flag column; |
outcome.name, event.name, treat.name |
Character. Column names in
|
effect_measure |
Character. Required for |
Details
For "continuous" and "count" this is an exact identity, not
agreement to Monte Carlo error – the frame is the super-population and the
arithmetic is the same compute_aor() dispatch the DGM used. A tolerance
there would hide a real defect.
This is a thin dispatch onto the same per-family effect used for every region; it introduces no arithmetic of its own.
Value
A named numeric vector, thetaDagger_H and thetaDagger_Hc.
The frame beta(Hhat) is scored on
Description
One entry point per outcome family, dispatching on outcome_type rather
than exposing a separate function per family. Defaults reproduce the
simulation modules this replaces exactly.
Usage
fs_build_eval_frame(
dgm,
outcome_type = c("survival", "binary", "continuous", "count"),
eval_seed = 20260628L,
analysis_time = 84,
cens_adjust = log(1.5),
n_eval = NULL
)
Arguments
dgm |
A DGM object carrying |
outcome_type |
Character. One of |
eval_seed |
Integer. Fixes the single realization on the full pool.
Ignored for |
analysis_time, cens_adjust |
Survival only. |
n_eval |
Defunct. Non- |
Value
A data frame: the evaluation population.
What each family returns
-
survival – the entire fixed super-population, every subject exactly once (
replace = FALSE), under the same randomized/censored analysis the trials run. -
binary – the same construction on the GLM simulator: every subject once, under one fixed treatment/outcome realization.
-
continuous / count –
dgm$df_superunchanged. No simulation occurs. The mean difference is collapsible, so the target is an exact finite mean over the super-population: the scoring frame is the population.eval_seedis accepted and ignored here so that generic harness code need not branch, and the target carries zero Monte Carlo error – uniformity of the call surface is not sameness of the object.
Rejected arguments
analysis_time, cens_adjust and n_eval are survival-only. Supplying any
of them on a non-survival path is an error rather than a silent no-op:
quietly ignoring an argument that means something on another path is how
conventions drift apart.
Calibrated declaration threshold and family-wise size of the p-star screen
Description
Post-hoc, opt-in, reported diagnostics of the consistency screen, computed
from multiplier draws the package already produces. The screen admits a
candidate g when its consistency rate, rounded to pconsistency.digits,
is at least p_star; on the resample path that is exactly the standardized
statistic T(g) = (beta_hat(g) - c_cons) / sigma_D(g) clearing the
effective cutoff z_pstar = qnorm((1 + pcons_eff) / 2) (see the Rounded
admission rule section). This function reports how often the maximum of
that statistic over the candidate family would clear the screen's cutoff
under the null perturbation law (fw_size), and the cutoff that would hold
the family-wise declaration rate at alpha (kappa_hat).
Usage
fs_declaration_calibration(
fit,
alpha = 0.05,
family = c("prereduction", "reduced"),
...,
c0 = NULL
)
## S3 method for class 'fs_declaration_calibration'
print(x, ...)
Arguments
fit |
A |
alpha |
Target family-wise declaration rate, in (0, 1). Default |
family |
|
... |
Only |
c0 |
|
x |
An |
Details
It reads a fitted object and returns a new one. It does not modify its
input, does not re-run the search, does not call the consistency engine,
and does not change what the search admitted: admitted_calibrated is the
set the calibrated rule would admit, reported beside admitted_current.
Value
An object of class fs_declaration_calibration: a list with
kappa_hat, fw_size, alpha, p_star, digits, digits_source,
pcons_eff, z_pstar (the effective z cutoff of the rounded screen),
consistency_method, consistency_method_source, c_cons,
c_screen, B, multiplier_law, quantile_type, family_source,
family_label, n_family_prereduction, n_family_reduced,
admitted_current (the candidates the executed screen admitted; NULL
when the fit is a bare fs_mr_inference() result), admitted_calibrated,
admitted_pstar (the relabelled current rule over the same family),
screened (the candidates the consistency screen evaluated, when known),
Mstar, beta_hat, sigma_D, T_hat, column_sd, zstar_mean,
field_cor (families of at most 8 candidates, else NULL), and a
reduction list recording the replay of the near-duplicate reduction.
When c0 is given it also carries c0, a list with table (one row per
c0: c0, c0_cmp, kappa_hat, fw_size, pstar_settable,
pstar_achievable, pcons_eff_settable, z_eff_settable, z_gap,
digits_fine, pstar_fine, z_gap_fine, n_admitted_calibrated, the 0.90 / 0.95 / 0.99 quantiles of Mstar_c0,
and is_c2), admitted_calibrated (a list indexed by c0), Mstar_c0
and source ("capture" or "field_matrix"). Every other element is
unchanged by c0.
Definitions
With shared multipliers xi[b, i] (one vector per draw b, reused across
every candidate) and the dfbeta influence db[g, i]:
-
Zstar[b, g] = sum_i xi[b, i] * db[g, i] / sigma_D(g), withsigma_D(g)^2 = sum_i db[g, i]^2– the robust scale the screen itself uses, never a model-based standard error. -
Mstar[b] = max_g Zstar[b, g]over the family: one-sided, in the harm direction. -
kappa_hat(alpha)is the empirical (type = 1)1 - alphaquantile ofMstar. -
fw_size = mean(Mstar > z_pstar),z_pstar = qnorm((1 + pcons_eff) / 2). Its threshold is the fit's ownp_starat the fit's ownpconsistency.digits, notalpha: it is the family-wise size of the screen as implemented, not of the calibrated rule. The calibrated admission rule is
beta_hat(g) >= max(c_screen, c_cons + kappa_hat * sigma_D(g)), the current rule withz_pstarreplaced bykappa_hat.
Rounded admission rule
The screen admits on round(Pcons, digits) >= p_star, with digits the
fit's pconsistency.digits. With p_star rounded up to the
10^-digits grid (g), that is Pcons >= pcons_eff = g - 0.5 * 10^-digits,
a bar below p_star itself (0.895 for p_star = 0.90, digits = 2). On
the resample path Pcons = 2 * pnorm(T) - 1, so the screen is
T >= z_pstar = qnorm((1 + pcons_eff) / 2). fw_size, admitted_pstar
and the settable p* are computed at that effective threshold; fw_size
is therefore the family-wise size of the screen as implemented and depends
on pconsistency.digits. kappa_hat does not: it is a quantile of the
maximum statistic and is compared with T directly, with no rounding.
digits is read from the fit's args_call_all$pconsistency.digits; when
absent (a bare fs_mr_inference() result, or an older fit) the
subgroup.consistency() default of 2 is used. digits_source records
which.
The correspondence with T holds only under
consistency_method = "resample". Under "split", Pcons is a split
proportion k / n_valid and admission is not a threshold on T, so
fw_size, z_pstar, admitted_pstar and the settable-p* columns are
NA / NULL rather than computed; kappa_hat and the calibrated rule are
still returned. The method is read from args_call_all$consistency_method;
a bare fs_mr_inference() result is treated as resample (the field is the
closed-form statistic), recorded in consistency_method_source.
Family
family = "prereduction" (the default, and the only value meant for the
calibrated rule) takes the maximum over the multiplier-resampling family:
every enumerated candidate of at most maxk factors meeting the size
minimum, before any reduction keyed on fitted quantities. The
near-duplicate reduction keys on sample-fitted summaries, so a family
reduced by it is outcome-dependent; the pre-reduction family is
covariate-measurable, and is also the conservative choice (a larger family
raises the maximum). The per-arm event minima are not replayed in that
family, which makes it a superset of the one the search chose among.
family = "reduced" removes the candidates the near-duplicate reduction
removed and is a diagnostic only, so the gap is measurable. Its result is
conditional on the realized family and is labelled so; it is never meant
for admission. It needs a forestsearch fit (the reduction lives in the
identifier) and the full field matrix (keep_field_matrix = TRUE).
Protected null level
c1 and c2 state the claim; c0 states what the claim is protected
against: a clinically specified, pre-specified benefit level (for example
HR 0.75). The family-wise declaration rate is then controlled at alpha
whenever every candidate's true effect is at least as good as c0, rather
than only when every candidate sits at c2. With
delta_g = (c_cons - c0_cmp) / sigma_D(g):
-
Mstar_c0[b] = max_g { Zstar[b, g] - delta_g }; -
kappa_hat(c0)is its empirical (type = 1)1 - alphaquantile; -
fw_size(c0) = mean(Mstar_c0 > z_pstar), at the rounded admission rule; admission is unchanged in form,
T(g) >= kappa_hat(c0);guidance on what to set to run
kappa_hat(c0)as ap*screen at the fit'sdigits:pstar_settable, the smallestp*on the10^-digitsgrid whose effective threshold is at or abovekappa_hat(c0); that threshold aspcons_eff_settable(Pcons scale) andz_eff_settable(z scale);z_gap = z_eff_settable - kappa_hat(c0)(positive = conservative); anddigits_fine/pstar_fine/z_gap_fine, the smallestdigits(searched over 1 to 12) at which the gap falls below 0.01, with itsp*. When nop* <= 1reacheskappa_hat(c0)at the fit'sdigits,pstar_achievableisFALSEand the settable columns areNA; the print method says so. This is guidance on what to set, not an identity: the settable screen iskappa_hat(c0)rounded up to what the grid can express, so it admits a subset of whatkappa_hat(c0)admits.
If every candidate's true effect is c0_cmp, T(g) is centred at
-delta_g, so the null law of max_g T(g) is that of the shifted maximum.
At c0 = c2, delta_g = 0 and every quantity is the unshifted one.
c0 is on the natural scale of the consistency threshold c2 (the HR on
the survival path; the ratio for OR / RR / IRR; the difference for RD /
MD) and is mapped exactly as c2 is: log() for ratio measures, identity
otherwise. c0 <= c2 is required. A non-positive c0 on a ratio path
(a log supplied by mistake) errors; on an identity path a wrong-scale
c0 cannot be detected.
The shifted maxima are read from the capture when the fit carries them
for every requested c0 (declaration_c0 at fit time); otherwise they are
computed from the stored field matrix (keep_field_matrix = TRUE);
otherwise the call errors. It never falls back to the unshifted maximum.
family = "reduced" always computes from the field matrix.
Inputs
The field is retained only on request. Fit with
mr_inference = TRUE and
mr_inference_args = list(keep_declaration_field = TRUE) (add
keep_field_matrix = TRUE for the reduced diagnostic), or call
fs_mr_inference(..., keep_declaration_field = TRUE) and pass its result.
The multiplier law and B are those of the MR draws and are recorded in
the result.
See Also
fs_mr_inference(), fs_family_report(), fs_fdr_report().
Examples
## Not run:
fit <- forestsearch(df, ..., mr_inference = TRUE,
mr_inference_args = list(keep_declaration_field = TRUE))
fs_declaration_calibration(fit, alpha = 0.05)
## End(Not run)
Design-time feasibility of a DGM's planted region
Description
Draws replicates through the DGM's own generator and reports, per sample
size, how often the planted harm region Q would be undeclarable
(too small for the search's own size test), how often it falls under the
per-arm events floor, and how often the analysis estimand does not exist
on it.
Usage
fs_dgm_feasibility(
dgm,
n = c(500, 750, 1000, 2000),
n.min = 60,
d0.min = 10,
d1.min = 10,
n_rep = 200L,
tolerance = 0.05,
effect_measure = NULL,
seed = NULL,
rand_ratio = 1
)
## S3 method for class 'fs_dgm_feasibility'
print(x, ...)
Arguments
dgm |
An object of class |
n |
Numeric vector of sample sizes to evaluate. |
n.min, d0.min, d1.min |
The search's own floors, passed so the report describes the campaign that will actually run. Defaults are the package defaults (60, 10, 10). Nothing is imposed: these are read, not applied. |
n_rep |
Integer. Replicates per sample size (default 200). |
tolerance |
Numeric. The largest undeclarable share that still counts as feasible (default 0.05). |
effect_measure |
Character or |
seed |
Integer or |
rand_ratio |
Numeric. Passed to |
x |
An |
... |
Ignored. |
Details
The motivating case: at n = 500 a region averaging 48.5 subjects is
at or below n.min = 60 in 96\
structurally unable to recover a region of the planted size – what it
declares is a larger overlapping region. That is a property of the design,
knowable before any replicate is run, and this function is how to know it.
Value
An object of class "fs_dgm_feasibility": a list with
$table (one row per n), $feasible (a single logical
– TRUE when every undeclarable share is at or below
tolerance), and $args.
What is counted
Per n, over n_rep draws:
share_undeclarableshare with
|Q| \len.min. The comparison is<=, exactly assubgroup.search()tests it (nx <= n.minrejects), so a region must exceedn.minto be declarable.share_under_eventsshare with fewer than
d0.minevents in the control arm ord1.minin the treated arm. Binary and survival only – the floor is skipped entirely for continuous and count outcomes.share_nonestimableshare on which the estimand does not exist, under the same per-estimand condition the estimator boundary applies: OR needs all four cells
\ge 1; RR, IRR and HR need an event in each arm; RD, IRD and MD have no condition and this share is always 0.
Together with per-arm cell summaries (mean, 5th and 95th percentile, and minimum of each cell across replicates).
Drawing through the DGM's own generator
Replicates are drawn with simulate_from_glm_dgm() – the same
function the campaign templates call – never a re-implementation, so the
population this reports on is the population a campaign would run on.
This function does not change the RNG kind. It seeds each replicate
with set.seed() under whatever generator is current and restores the
caller's RNG state on exit. The caller's own obligation is the mirror of
that one: build and calibrate the DGM before any replicate switches the
RNG kind. Calibrating a DGM after a switch to "L'Ecuyer-CMRG"
yields a different super-population, and every count here would then
describe a population no campaign ever used.
Note
Survival DGMs are not supported, and what they would need is
this: the OC family's contract is inherits(dgm, "glm_dgm") with
a df_super carrying flag_harm. setup_gbsg_dgm()
produces a different object whose own generator is
simulate_from_dgm(dgm, n, analysis_time, cens_adjust, seed) –
a different signature, and it takes two censoring arguments that have no
GLM counterpart and that change the per-arm event counts this function
reports. Supporting it means a generator-dispatch layer plus passing
analysis_time and cens_adjust through; that is a separate
change and is deliberately not guessed at here.
See Also
simulate_from_glm_dgm, fs_oc_predict
Examples
set.seed(1)
N <- 400
d <- data.frame(treat = rbinom(N, 1, 0.5),
age = round(rnorm(N, 55, 12)),
sex = factor(rbinom(N, 1, 0.5), levels = 0:1))
d$y <- rbinom(N, 1, plogis(-0.5 + 0.3 * d$treat))
dgm <- generate_glm_dgm(data = d, factor_vars = "sex",
continuous_vars = "age", outcome_var = "y",
treatment_var = "treat", outcome_type = "binary",
effect_measure = "OR", subgroup_vars = "age",
subgroup_cuts = list(age = list(type = "greater",
quantile = 0.55)),
k_inter = 1, n_super = 2000, seed = 8316951)
feas <- fs_dgm_feasibility(dgm, n = c(500, 1000), n_rep = 50)
feas
feas$feasible
Sampling Scale of the Difference-in-Means Estimator
Description
Computes the finite-population sampling scale of the within-region difference-in-means estimator directly from a data generating mechanism's individual-level potential outcomes.
Usage
fs_dgm_scale(
dgm,
regions = NULL,
harm_col = "flag_harm",
rand_ratio = 1,
labels = c("Q", "Qc", "S")
)
Arguments
dgm |
An object of class |
regions |
Optional named list of region specifications. Each element is
either a logical vector of length |
harm_col |
Character. Name of the true-region indicator column in
|
rand_ratio |
Numeric. Treatment:control randomisation ratio, so the
randomisation probability is |
labels |
Optional character vector of length 3 renaming the default
regions. Ignored when |
Details
For a region g the estimator is the within-region difference in arm
means. Conditioning on the realized arm counts,
\mathrm{Var}[\hat\beta(g)] = V_{\mathrm{eff}}(g) / (n P(g)),
with the region- and n-free constant
V_{\mathrm{eff}}(g) = V_1(g)/p + V_0(g)/(1-p),
where p is the randomisation probability and V_w(g) is the total
within-arm outcome variance in g under arm w:
V_w(g) = m_g[v_w] + V_g[\mu_w].
Here \mu_w is the arm mean surface, v_w(x) = \mathrm{Var}(Y \mid
X = x, W = w) the conditional outcome variance, and m_g[\cdot],
V_g[\cdot] finite-population moments over the rows of the
super-population lying in g. The conditional variance is supplied by
the outcome family:
- continuous
v_w = \sigma^2, the residual variance (constant).- binary
v_w = p_w (1 - p_w).- count
v_w = \mu_w(Poisson).
Value
An object of class c("fs_dgm_scale", "list"):
regionsData frame, one row per region, with columns
region,n_g,P_g,m_mu0,m_mu1,m_tau,V_mu0,V_mu1,V_tau,C_mu0_tau,v_cond0,v_cond1,V_arm0,V_arm1,bracket,V_eff.sigmaResidual standard deviation for continuous outcomes, otherwise
NA_real_.outcome_type,effect_measureCopied from
dgm.rand_ratio,p_treatAllocation used.
n_superSuper-population size.
Balanced-arm reading
At 1:1 allocation V_{\mathrm{eff}}(g) = 2\{V_0(g) + V_1(g)\}, and
bracket = V_{\mathrm{eff}}(g)/4 admits the decomposition
\sigma^2 + V_g[\mu_0] + C_g[\mu_0, \tau] + \tfrac12 V_g[\tau],
with \tau = \mu_1 - \mu_0 the conditional average treatment effect.
The component columns are returned for every region and every outcome type;
the decomposition above is exact for continuous outcomes at 1:1 allocation.
Away from 1:1, bracket remains V_{\mathrm{eff}}/4 but no longer
equals that sum, so prefer V_eff for any downstream calculation.
Idealisations
V_eff treats the region size and the arm split as fixed at their
expectations. A trial draws n_g \sim \mathrm{Bin}(n, P(g)) and
n_1 \mid n_g \sim \mathrm{Bin}(n_g, p), and by Jensen's inequality
E[1/n_1 + 1/n_0] and E[1/n_g] exceed their fixed-count
counterparts. Use fs_scale_se with jensen = TRUE for
the unconditional standard deviation.
Scope
Exact for identity-scale effect measures ("MD", "RD",
"IRD"), where the estimator is a difference in means. Ratio measures
("OR", "RR", "IRR") require a delta-method layer and are
rejected rather than silently approximated.
See Also
Examples
## Not run:
dgm <- generate_glm_dgm(..., outcome_type = "continuous",
effect_measure = "MD")
sc <- fs_dgm_scale(dgm)
sc$regions[, c("region", "P_g", "bracket", "V_eff")]
# Standard deviation of the estimator on the true region at n = 500
fs_scale_se(sc, n = 500, region = "Q")
# Arbitrary regions, e.g. a candidate family
fs_dgm_scale(dgm, regions = list(big = dgm$df_super$age > 40))
## End(Not run)
Materialize the maxeff candidate family as 0/1 membership columns
Description
Reconstructs, for every candidate in the deduplicated maxeff family kept on
a forestsearch(sg_focus = "maxeff") fit, its per-subject 0/1 membership on
the estimation frame fit$df.est. Each candidate's cut labels (e.g.
"{er <= 0}", or a negated cut "!{size <= 20}") are mapped back to the
aligned dummy columns of fit$df.est via fit$confounders.evaluated /
fit$confounders.candidate, and a candidate's membership is the
intersection (AND) over its cuts. The result is the materialized family
that guohe_algorithm3() consumes; fs_to_guohe() calls this internally.
Usage
fs_export_maxeff_family(fit, prefix = "sg_")
Arguments
fit |
A |
prefix |
Character prefix for the generated membership column names
( |
Value
A list with elements data (fit$df.est with one appended 0/1
membership column per candidate), candidates (the appended column
names), selected (the column name of the forestsearch-selected
candidate; the family table is sorted selected-first), and family (the
candidate family table fit$grp.consistency$out_sg$result).
See Also
fs_to_guohe(), fs_assert_membership()
Examples
set.seed(11)
n <- 200
df <- data.frame(id = seq_len(n), treat = rbinom(n, 1, 0.5),
age = round(rnorm(n, 55, 12)),
bm = factor(rbinom(n, 1, 0.5)))
harm <- df$age <= 50 & df$bm == "1"
tt <- rexp(n, 0.05 * exp(log(2.5) * df$treat * harm))
df$time <- pmin(tt, 60)
df$event <- as.integer(tt <= 60)
fit <- forestsearch(
df.analysis = df,
outcome.name = "time", event.name = "event",
treat.name = "treat", id.name = "id",
confounders.name = c("age", "bm"),
sg_focus = "maxeff", use_grf = FALSE, use_lasso = FALSE,
fs.splits = 50, n.min = 30, d0.min = 5, d1.min = 5,
hr.threshold = 1.1, hr.consistency = 1.0,
pconsistency.threshold = 0.5, maxk = 2,
details = FALSE, plot.sg = FALSE,
parallel_args = list(plan = "sequential", workers = 1L,
show_message = FALSE)
)
exp_fam <- fs_export_maxeff_family(fit)
head(exp_fam$candidates)
exp_fam$selected
Report which stages of the candidate family are data-dependent
Description
A user who sets use_lasso = FALSE, use_grf = FALSE,
use_dina = FALSE and vi.grf.min = NULL may reasonably believe
the candidate family forestsearch searches is now fixed. It
is not. Continuous cuts are placed at sample quantiles
(get_FSdata()); the prevalence, redundancy and size floors
(minp, rmin, n.min) are applied to sample
counts (subgroup.search()); and the consistency stage removes
near-duplicate candidates keyed on their fitted statistics, with no
argument that turns it off (remove_near_duplicate_subgroups()).
Those stages are the method, not screening. This function says so, for a
given argument set, in a form that cannot overpromise.
Usage
fs_family_report(x, data = NULL, outcome_type = NULL)
## S3 method for class 'fs_family_report'
print(x, ...)
Arguments
x |
Either a named list of |
data |
Optional data frame. When supplied, the report is grounded in
counts: the number of cut columns |
outcome_type |
One of |
... |
Ignored. |
Details
It reports and changes nothing. It never fits a model, never runs the search, and is called by nothing else in the package. No combination of arguments makes the family deterministic while cuts are placed at sample quantiles; the best a caller can do is switch off the disableable stages, and the printed footer lists what remains.
The report mirrors forestsearch()'s own argument resolution rather
than reading arguments at face value: sg_focus aliases are
normalised; sg_focus = "maxeff" zeroes minp, rmin and
pconsistency.threshold, sets stop_threshold to NULL,
use_twostage to FALSE and max_subgroups_search to
Inf, and disables the effect floor; stop_threshold is reset
to NULL for every focus other than "maxeffCons" (it is
meaningful only there); max_n_confounders is applied only inside the
GRF variable-importance block, so it is inert when vi.grf.min is
NULL; and the per-arm event floors d0.min / d1.min are
skipped for continuous and count outcomes.
Value
A data frame of class c("fs_family_report", "data.frame")
with one row per stage and columns stage, arguments (the
governing forestsearch() formals), values (their resolved
values, formatted), status (one of "deterministic",
"disabled", "inert", "data-dependent",
"data-dependent (not disableable)") and note (one line:
why, and what would change it; NA where nothing would).
Attributes: verdict (one sentence), status_counts (a
named integer vector), data_supplied (logical), and – when
data is supplied – n_cut_columns and
n_combinations.
See Also
forestsearch, fs_oc_family_enumerate
(the population-frame enumeration used by the operating-characteristic
wrapper, which is deterministic in the DGM precisely because it does not
cut at sample quantiles).
Examples
rep <- fs_family_report(
list(confounders.name = c("age", "preanti", "hemo"),
conf.cont_jcuts = list(age = 10, preanti = 10),
n.min = 60, maxk = 2, sg_focus = "maxeffCons",
use_lasso = FALSE, use_grf = FALSE, use_dina = FALSE,
vi.grf.min = NULL),
outcome_type = "continuous")
print(rep)
attr(rep, "verdict")
# grounded in counts: cut columns and combinations on a data frame
set.seed(1)
df <- data.frame(age = round(rnorm(200, 40, 8)),
preanti = round(rexp(200, 1 / 500)),
hemo = rbinom(200, 1L, 0.1))
rep2 <- fs_family_report(
list(confounders.name = c("age", "preanti", "hemo"),
conf.cont_jcuts = list(age = 10, preanti = 10), maxk = 2,
use_lasso = FALSE, use_grf = FALSE, vi.grf.min = NULL),
data = df, outcome_type = "continuous")
attr(rep2, "n_cut_columns"); attr(rep2, "n_combinations")
Operating False-Discovery Rate of a Fitted ForestSearch Analysis
Description
Characterizes the operating false-discovery rate (FDR) of a fitted
forestsearch analysis and tests whether requiring harm
confirmation on the de-biased effect reduces it. This is a measurement
tool, not a calibration tool:
it does not solve for thresholds that control an FDR target. For the
(c_1, c_2) pair the analysis actually used, it reports two quantities
under a null:
Usage
fs_fdr_report(
fs_analysis,
df_analysis,
null = c("hom", "harm"),
c_confirm = c(0.9, 1, 1.25),
nsims = 1000L,
n = NULL,
confounders = NULL,
n_super = 25000L,
analysis_time = 84,
cens_adjust = log(1.5),
event_min = 5L,
n_workers = NULL,
seed_base = 8316951L,
quiet = FALSE
)
Arguments
fs_analysis |
A fitted |
df_analysis |
The analyst's data frame. Used only to default the
simulated sample size ( |
null |
Character vector; which null(s) to simulate. |
c_confirm |
Numeric vector of de-biased-HR harm-confirmation
thresholds at which to report B, i.e. the values |
nsims |
Integer; null replicates per null. Default 1000. |
n |
Integer or |
confounders |
Character vector or |
n_super |
Integer; super-population size for the DGM. Default 25000. |
analysis_time |
Numeric; administrative censoring time (months). Default 84. |
cens_adjust |
Numeric; censoring shift passed to
|
event_min |
Integer; minimum events for a replicate to be scored. Replicates below this (or with a single treatment arm) are counted as visible failures, never silently dropped. Default 5. |
n_workers |
Integer or |
seed_base |
Integer; base seed. Each null uses a disjoint seed block and each replicate an explicit seed, so results are plan-invariant. Default 8316951. |
quiet |
Logical; suppress progress and per-null messages. Default
|
Details
-
A (raw FDR):
A = P(\text{declare any harm subgroup} \mid c_1, c_2). -
B(g) (confirmed FDR):
B(g) = P(\text{declare AND de-biased } HR \ge g). Because harm confirmation can only remove declarations,B(g) \le Aalways; the gapA - B(g)is what the de-biasing step buys at confirmation thresholdg.
All thresholds and identifier knobs are inherited from
fs_analysis$args_call_all, so the FDR reported is that of the analysis
as run. Multiplier resampling is forced on internally
(mr_inference = TRUE) regardless of the fitted setting, because B
requires the per-replicate de-biased HR.
What "false discovery" means here
Under "hom" the treatment effect is real but uniform, so any declared
subgroup is a manufactured heterogeneity – the winner's-curse-of-search
rate. Under "harm" a real modifier exists and the flagged region
sits exactly at HR = 1.0, so a small upward fluctuation reads as harm; this
is the more adversarial null. The declared cut need not overlap
flag_harm: under either null any declaration is a false discovery.
The confirmation accounting (why n_deb_na matters)
B(g) counts a declaration only when its de-biased HR is finite and clears
g; a declaration whose de-biased HR could not be computed counts as
unconfirmed. Such cases are counted in n_deb_na and reported, so
that A - B is never mistaken for pure de-biasing benefit when part
of it is de-biased-HR unavailability.
Inheritance and rebinding
Identifier knobs (thresholds, subgroup_method, consistency_method,
conf_force, conf.cont_jcuts, maxk, MR arguments, ...)
transfer verbatim from args_call_all. The data-binding arguments –
the five column names and the confounder set – are rebound to the DGM,
because simulate_from_dgm() emits fixed names
(y_sim / event_sim / treat_sim / id / flag_harm). mr_inference
is forced TRUE so B is computable.
Value
An object of class "fs_fdr_report": a list with
fdrData frame, one row per null x
c_confirm, withA,A_lo,A_hi(raw FDR and 95% Wilson CI),B,B_lo,B_hi(confirmed FDR and CI),reduction(A - B), and the accounting columnsn_valid,n_declared,n_deb_na,fs_errors,unhandled.hr_structureData frame of the realized hazard ratios each DGM encodes – the structural check that the nulls are different worlds.
verdict"CLEAN"if no replicate failed across any null, else"WARNING".n_errorsTotal failed replicates (forestsearch errors + unhandled).
c1,c2,n,nsims,nullsThe inherited thresholds and run settings.
call_argsReproducibility metadata.
See Also
forestsearch (and its vocabulary section),
fs_mr_inference for the MR step this forces on,
mr_estimates_table,
forestsearch_bootstrap_dofuture for the full bootstrap,
fpr_calibration,
setup_gbsg_dgm, calibrate_k_inter,
simulate_from_dgm
Examples
## Not run:
# A fitted survival/Cox analysis (thresholds and knobs live on the object)
fit <- forestsearch(
df.analysis = gbsg_df,
outcome.name = "time", event.name = "status",
treat.name = "treat", id.name = "id",
confounders.name = c("er", "age", "meno", "pgr", "nodes", "size", "grade"),
is.RCT = TRUE, est.scale = "hr",
subgroup_method = "consistency", consistency_method = "resample",
hr.threshold = 1.25, hr.consistency = 1.0,
pconsistency.threshold = 0.90, maxk = 2,
mr_inference = TRUE)
# Characterize its operating FDR and whether harm confirmation helps
rep <- fs_fdr_report(fit, gbsg_df, nsims = 1000, n_workers = 100)
rep # prints the realized-HR table and the A/B table
rep$fdr # one row per null x c_confirm
# Reproduce the published reference (n = 700): A ~ 0.138 (hom), 0.278 (harm)
fs_fdr_report(fit, gbsg_df, n = 700, nsims = 1000, n_workers = 100)
## End(Not run)
Stem tag for a (subgroup_method, sg_focus) pair
Description
Returns the short label used in result-file stems and plot annotations for
the rule that a given (subgroup_method, sg_focus) combination
actually runs – not the spelling the caller passed. The two differ
whenever an alias or a behavioural synonym is in play, which is precisely
when a hand-written tag goes wrong.
Usage
fs_focus_tag(subgroup_method, sg_focus)
Arguments
subgroup_method |
Character scalar. One of |
sg_focus |
Character scalar. Any accepted |
Details
The collapse is engine-specific:
"consistency"eff,hrandmaxconsall tag as"maxcons"– they name the consistency argmax.maxeffandmaxeffConsstay distinct, because the consistency floor genuinely separates them here."dina","grf"eff,hr,maxcons,maxeffandmaxeffConsall tag as"eff". Neither engine computes a Pcons, so the consistency qualifier has nothing to bind to and all five rank byorder(-eff).
The band foci collapse the same way on every engine: hrMaxSG tags as
"effMaxSG" and hrMinSG as "effMinSG". maxSG and
minSG are canonical foci in their own right and pass through.
Value
Character scalar; the stem tag.
See Also
forestsearch for the per-engine semantics of each
focus.
Examples
fs_focus_tag("consistency", "eff") # "maxcons"
fs_focus_tag("dina", "eff") # "eff"
fs_focus_tag("dina", "maxeffCons") # "eff" -- not "maxeffCons"
fs_focus_tag("grf", "hrMaxSG") # "effMaxSG"
Identification structure for one set of replicates
Description
Classifies every detected replicate's realized rule against a planted
anchor / partner / proxy triple, and returns covariate involvement,
the structure composition, and classification accuracy.
Usage
fs_identification_structure(
results,
anchor,
partner,
proxy,
confounders = NULL,
sg_def_col = "sg_def",
covs_col = "covs"
)
Arguments
results |
A per-replicate results data frame, a bundle carrying
|
anchor |
Character. The covariate the search is expected to recover. |
partner |
Character. The true second covariate of the planted region. |
proxy |
Character. A correlated covariate that may substitute for
|
confounders |
Character vector. The analysis covariate pool, for the involvement vector. Defaults to the three named roles. |
sg_def_col, covs_col |
Character. Columns holding the realized rule and its projected covariate names. |
Value
A list with det_rate, n_detected, n_total, involvement,
structure (a named share vector), and the four accuracy means.
The classification
Each detected replicate falls in exactly one of five mutually exclusive
categories, tested in this order: the anchor is absent (anchor missed);
the rule names the anchor and nothing else (anchor only); the rule pairs
the anchor with the true partner; the rule pairs it with the proxy; or
the rule pairs it with something else. The order matters – a rule naming
both the partner and the proxy counts as the true pairing.
Identification structure across a sweep
Description
Applies fs_identification_structure() across a named collection of
bundles – typically one per trial size – and returns a single object the
plot and table helpers consume. This is the package-level replacement for
the inline fs_identification_figures*.qmd fragments.
Usage
fs_identification_summary(
bundles,
anchor,
partner,
proxy,
confounders = NULL,
sweep_label = "Trial size (n)",
sg_def_col = "sg_def",
covs_col = "covs"
)
Arguments
bundles |
A named list. Each element is a results data frame, a bundle
carrying |
anchor, partner, proxy |
Character. The planted-structure triple. |
confounders |
Character vector. The analysis covariate pool. |
sweep_label |
Character. Axis label for the sweep dimension. |
sg_def_col, covs_col |
Character. Passed through. |
Value
An object of class fs_identification: a list with per_point
(the raw aggregates), tidy involvement and structure data frames, an
accuracy data frame, and the triple it was built with.
Examples
## Not run:
id <- fs_identification_summary(
list(`500` = "results/fs_..._n500_combined_1_500.rds",
`1000` = "results/fs_..._n1000_combined_1_500.rds"),
anchor = "er", partner = "meno", proxy = "age",
confounders = c("er", "age", "meno", "pgr", "nodes", "size", "grade"))
plot_fs_identification_involvement(id)
plot_fs_identification_structure(id)
fs_identification_table(id)
## End(Not run)
Identification-structure table
Description
The supplement's structure table: detection rate, the five structure shares,
anchor recovery, and the conditional partner-versus-proxy split, one column
per sweep point. Returned as a data frame by default so a LaTeX document can
style it however it likes; as_gt = TRUE renders it with gt.
Usage
fs_identification_table(x, as_gt = FALSE, digits = 1)
Arguments
x |
An |
as_gt |
Logical. Render with gt instead of returning the data frame. |
digits |
Integer. Digits on the percent labels. |
Value
A data frame, or a gt_tbl. The accuracy sentence the supplement
carries as a table footnote is attached as the "accuracy_note"
attribute of the data frame.
Uniform (kappa) calibration of the field interval
Description
Computes, at analysis time and from the trial's own influence structure,
the smallest widening factor kappa on kappa_grid such that the two-sided
field interval c + kappa * (q - c) attains 1 - alpha coverage
uniformly over a winner-profile protection family: hypothetical true
fields equal to the gate's shrunk field w with the winner's entry set to
max_{g != winner} w_g + delta * sigma_sel, for delta on delta_grid
(in SE units of the winner). Nothing is tuned to any simulation design.
Usage
fs_mr_field_uniform(
B = NULL,
Sigma = NULL,
w,
sel,
sigma_sel,
p_hat,
t_g = NULL,
reselection = "maxeff",
sz = NULL,
effect_neighborhood = 0.1,
selection_rule = "neighborhood",
log_scale = TRUE,
sdv = NULL,
zcons_c = NULL,
delta_grid = seq(0, 4, by = 0.5),
mass = 0.99,
M_cap = 40L,
R_rep = 300L,
R_out = 300L,
R_in = 150L,
alpha = 0.05,
kappa_grid = seq(1, 3, by = 0.01),
seed = NULL
)
Arguments
B |
Optional n x K influence matrix (the gate's |
Sigma |
Optional K x K covariance, used when |
w |
Numeric K: the gate's shrunk field ( |
sel |
Integer: the winner's index in |
sigma_sel |
The winner's SE on the working scale ( |
p_hat |
Numeric K: re-selection frequencies from the gate's
multiplier pass ( |
t_g |
Numeric K admission floors, or |
reselection, sz, effect_neighborhood, selection_rule, log_scale |
The
gate's selection-map configuration; |
sdv |
Numeric K per-candidate SEs (needed only for |
zcons_c |
Consistency centre |
delta_grid |
Winner-separation grid in SE units (H2 default 0-4 by 0.5). |
mass, M_cap |
Mass-carrying reduction controls (H3 as adjusted at the Gate 1 adjudication, 2026-09-05: cap 40, so the >= 0.99 mass target is reachable on enumerated forestsearch families; original default was 12). |
R_rep, R_out, R_in |
Monte Carlo sizes per delta (H4 defaults 300/300/150). |
alpha |
Two-sided miscoverage target (0.05). |
kappa_grid |
Candidate widening factors (default 1 to 3 by 0.01; the grid was extended past 2 at the Gate 1 adjudication after ceiling hits). |
seed |
Optional integer; drawn once at entry. |
Details
The computation restricts to a mass-carrying candidate set (the smallest
M <= M_cap candidates by re-selection frequency whose cumulative share
of the re-selection mass is at least mass, the winner always included),
then per delta replays R_rep hypothetical trials with Gaussian
multipliers of covariance exactly Sigma_hat (via crossprod(B, xi), or
a Cholesky root of a supplied Sigma), applies the gate's own selection
map (thresholded argmax under reselection = "maxeff", vectorized; any
other rule via the per-draw .fs_mr_select path), drops trials with no
winner (conditioning on detection, as in the real analysis), and runs the
field procedure on each hypothetical trial to obtain the coverage
profiles C1(delta) (one-sided) and C2(delta; kappa) (two-sided, widened).
The guarantee, as documented for the gate: the one-sided bound is
uniformly valid; the two-sided interval widened by the returned kappa
is uniformly valid over the winner-profile family; the plain quantile
interval is approximate.
Value
List: kappa (kappa*), kappa_mcse (Monte Carlo SE of kappa*
from the binding profile's binomial error and the local slope of
min-delta C2 in kappa), M, mass_covered, keep (indices), minC1,
C1 (profile over delta_grid), C2_k1 (profile at kappa = 1),
C2_kstar (profile at kappa*), kappa_grid, C2_min (min-over-delta
profile over kappa_grid), n_kept (trials with a winner per delta),
delta_grid, timing_seconds.
Multiplier resampling (MR) for a selected forestsearch subgroup
Description
Computes a multiplier-bootstrap approximation of the bootstrap
bias-corrected treatment effect for the selected subgroup and flags whether
it is still consistent with harm. Reuses the resample consistency engine's
per-candidate treatment dfbeta; see the file header for the closed form.
Usage
fs_mr_inference(
df,
candidates,
spec,
selected_members,
admission,
t_confirm = NULL,
confirm_rule = c("point", "ci"),
reselection = c("maxcons", "maxeff", "maxSG", "minSG", "effMaxSG", "effMinSG"),
effect_neighborhood = 0.1,
selection_rule = c("neighborhood", "pareto", "both"),
draws = 2000L,
multiplier = c("poisson", "gaussian", "rademacher"),
include_complement = FALSE,
ci_method = c("field", "ij", "wald"),
seed = NULL,
return_reselection = TRUE,
field_R_out = 1000L,
field_R_in = 500L,
field_uniform = FALSE,
field_M_cap = NULL,
field_complement = TRUE,
field_decompose = FALSE,
field_scale_complement = c("selected", "none"),
ij_residual = c("two_term", "winner", "winner_floor"),
field_recovery = FALSE,
keep_declaration_field = FALSE,
keep_field_matrix = FALSE,
declaration_c0 = NULL,
pconsistency.digits = NULL
)
Arguments
df |
Analysis data frame (the standardized |
candidates |
Named list of integer row-index vectors, one per screened candidate subgroup (the family the selection rule chose among). |
spec |
List with |
selected_members |
Integer row indices of the observed selected subgroup
( |
admission |
The resolved admission set, as returned by
This replaces the former
|
t_confirm |
Harm-confirmation threshold on the effect scale (HR/OR/
RR/IRR, or RD/MD for differences – not the working log scale). |
confirm_rule |
Which harm-confirmation rule to apply to the de-biased
estimate: |
reselection |
Bootstrap re-selection rule for the bias term; default
|
effect_neighborhood |
Band for the |
draws, multiplier, seed |
Multiplier-bootstrap controls. |
include_complement |
Logical. When |
ci_method |
|
return_reselection |
Logical. When |
field_R_out, field_R_in |
Outer and inner Monte Carlo sizes for
|
field_uniform |
Logical (default |
field_M_cap |
Optional override of the uniform sweep's mass-carrying
cap ( |
field_complement |
Logical (default |
field_decompose |
Logical (default |
field_scale_complement |
|
ij_residual |
Which IJ residual populates the reported de-biased
SE and interval ( |
field_recovery |
Logical (default |
keep_declaration_field |
Logical (default |
keep_field_matrix |
Logical (default |
declaration_c0 |
|
pconsistency.digits |
Integer or |
Details
This is post-selection inference on a completed analysis: it runs after
forestsearch() has chosen the subgroup and cannot change that choice. It
is the fast, refit-free counterpart of the full bootstrap (FB,
forestsearch_bootstrap_dofuture()), which re-runs the entire search in
every replicate; MR perturbs the influence contributions of a single fit
instead. MR approximates FB to leading order and does not replace it.
Value
List with the selected index/label, naive and debiased estimates
(effect scale, with approximate 95% CIs), selection_bias, fixed_bias,
selection_rate, mean_r, mean_r_c, the settings actually used (t_confirm,
confirm_rule, reselection, selection_rule, multiplier, draws),
harm_flag, family/subgroup sizes, and timing_seconds. The debiased
element carries se_ij, se_wald, var_ij, and ij_source; its CI uses
the IJ SE under ci_method = "field" (the default) and "ij" (the FB
analogue in both cases) and the robust SE under "wald".
When include_complement = TRUE, a complement element carries the
complement subgroup's naive/debiased estimates and bias terms in the
same form, including its own IJ variance. The complement's de-biased CI
follows the same SE convention as the winner's: the IJ SE under
ci_method = "ij" and "field" (under "field" the complement, like the
debiased element, is identical to the "ij" output), the robust SE
under "wald".
mean_r is the mean of the IJ residual r_b over ok_H – exactly
the draws entering the variance, not all draws, since on an excluded
draw r_b is not a meaningful quantity. The invariant is that it is
zero by construction whenever both bias terms share a denominator, which
is the convention the package implements: selection_bias and
fixed_bias both average over the draws that produced a winner, and the
IJ runs on that same set. A non-zero value therefore means the two terms
are being normalised differently somewhere – the defect corrected in
dad0415, whose signature was precisely a non-zero residual mean.
mean_r_c is the same quantity for the complement over use_c, and is
NA_real_ when no complement was fit. Both are diagnostics of the
correction's internal consistency rather than properties of the estimate,
which is why they sit beside selection_rate and not inside debiased.
Note the mixed scales in the flat debiased list: est/lower/upper/
lower_1s are on the effect scale, while se/se_ij/se_wald/
var_ij and the two bias terms are on the working (log, for ratio
measures) scale.
When the selected subgroup cannot be fit in the reconstructed family the
return is a short variant carrying only selected_index (NA),
selected_label, harm_flag (NA), settings, note and n_family;
consumers must tolerate that shape.
Under return_reselection = TRUE the full return additionally carries a
reselection element (see that argument); the short variant never does.
Under ci_method = "field" the full return additionally carries a
field element: lambda_mean, lambda_sd/se_field and the
Lambda* quantiles q05/q25/q50/q75/q95/q025/q975 on the working
scale; n_out_used, n_in_used_mean, R_out, R_in, seed_offset,
timing_seconds; and, on the effect scale (the debiased
convention), the second-order point estimate est2 (de-biased estimate
minus lambda_mean), the primary one-sided 95% bound lower_1s
(de-biased minus q95), the two-sided quantile interval
lower_2s/upper_2s, and the supplementary SE-type interval
lower_se/upper_se around est2.
Under field_recovery = TRUE the field element additionally carries
recovery: sens_H (the primary quantity – the mean share of the
identified patients that the re-selections retain), ppv_H, sens_Hc,
ppv_Hc and its alias npv_Hc; the containment quantiles q10, q50,
q90 and share_equal_1 (the share of draws whose re-selection contains
all of the observed subgroup); and the accounting n_draws, n_used,
n_skipped, n_selected, n_all, timing_seconds. Descriptive only;
absent when field_recovery = FALSE.
Under field_complement = TRUE (with include_complement = TRUE) the
field element additionally carries complement, the complement's own
field block in the same form and scales, inverted around the complement's
two-term de-biased estimate: lambda_mean, lambda_sd/se_field, the
seven Lambda*c quantiles; on the effect scale est2, the primary
one-sided 95% upper bound upper_1s (de-biased minus q05), the
one-sided lower bound lower_1s, the two-sided lower_2s/upper_2s
and the SE-type lower_se/upper_se; the draw accounting
n_out_used, n_in_used_mean, n_out_dropped_unfit; the fit accounting
n_complement_fits (distinct complement fits, the multiplier stage's
included), n_new_fits, share_draws_new_fit (share of outer + inner
draw-winner readings whose candidate the multiplier stage had not fit);
R_out, R_in, timing_seconds. A note replaces the numbers when
the selected complement is unfit or fewer than 2 outer draws survive.
With the complement field on, field also carries joint: the
simultaneous (harm lower, complement upper) pair from the aligned outer
draws (\Lambda^*_r, \Lambda^{*c}_r) – same multipliers, same
winners, no new draws. gamma is the equal-tail level (grid from
alpha down to alpha / 2, step 0.001; the largest value whose joint
probability P^*(\Lambda^* \le q_{1-\gamma}, \Lambda^{*c} \ge
q_\gamma) \ge 1 - \alpha), joint_prob the achieved probability,
lower_H/upper_Hc the calibrated pair (effect scale),
bonf_lower_H/bonf_upper_Hc the Bonferroni pair at gamma = alpha/2
with its bonf_joint_prob, corr the draws' correlation,
n_joint_draws, and the grid_gamma/grid_joint_prob profile. The
marginal one-sided bounds are unchanged. The top-level ij_residual
records the residual that populated the reported IJ SEs.
Alignment is assumed, not checked here
This is an engine-level entry point. Its arguments are a candidate family,
a specification, and a selected membership vector – the identifier
configuration that produced them is not visible, so the alignment
conditions MR requires (selection ranking on the inferential coefficient
\hat\beta(g), and a fixed candidate family) cannot be verified at
this level and are assumed to have been established upstream.
They are enforced by .validate_mr_configuration() at the three
configuration-visible entry points – forestsearch() under
mr_inference = TRUE, and forestsearch_bootstrap_dofuture() and
forestsearch_Kfold() under mr_in_replicates = TRUE. Calling
fs_mr_inference() directly bypasses those guards: a family ranked on
DINA's native tau-hat, on GRF's doubly-robust score, or on a GRF
policy-tree objective will still produce numbers, but they do not de-bias
the reported effect.
See Also
forestsearch() for the mr_inference switch and the
vocabulary section; mr_estimates_table() to render the result;
forestsearch_bootstrap_dofuture() for the full bootstrap (FB) this
approximates; fs_fdr_report() which sweeps c_confirm thresholds.
Operating-Characteristics Summary of a Multiplier-Resampling Payload
Description
Summarises a simulation payload produced by the multiplier-resampling (MR) harness into the estimation, coverage and classification quantities used in simulation summaries, returning plain numbers rather than a formatted table.
Usage
fs_mr_oc_summary(
payload,
estimators = NULL,
blocks = c("H", "Hc"),
digits = NULL
)
Arguments
payload |
A list with |
estimators |
Character vector selecting estimators, from
|
blocks |
Character vector of blocks, from |
digits |
Integer or |
Value
An object of class c("fs_mr_oc", "list"):
estimationData frame, one row per block and estimator:
n,avg,sd_emp,se_hat, the fourbias_*, the fourcov_*, andwidth.identificationData frame with one row per convention: detection rate, mean subgroup size, and mean sens/spec/ppv/npv.
targetsThe oriented target values used.
metaKey run metadata, echoed.
Orientation
Estimator columns (or_, nv_, fb_, mr_) are stored
on the ORIENTED scale, where positive means harm. betaHhat_H and
betaHhat_Hc are stored on the RAW outcome scale. The bridge is
orient = if (adverse_outcome) 1 else -1, applied here as the
simulation harness applies it. Getting this backwards silently negates every
bias.
Four targets, deliberately
Every estimator is scored against every target in both blocks, so that disagreement is visible rather than assumed away:
betaPer-replicate exact
\beta(\hat H)– the conditional estimand at the realised region.oraclePer-replicate refit on the true region in that trial.
structScalar structural potential-outcome effect in the true region; zero Monte Carlo error.
margScalar fitted effect in the true region on one realised draw; carries sampling noise.
Bias is ABSOLUTE, in outcome units. On an identity scale a complement effect can sit near zero, so a relative bias would explode.
Two classification conventions
Both are returned, because the project contains both and they disagree whenever detection is not near certain:
conditionalAveraged over DETECTED replicates only. This is the detected-only summary.
unconditionalAveraged over ALL replicates, following
build_classification_table()and Leon et al. (2024, Table 1), with non-detection scored as sens = 0, ppv = 0, spec = 1, npv = 1.
With detection at 999/1000 the two agree to three decimals; in a null cell at 0.930 they do not. Quote which one you mean.
See Also
Examples
## Not run:
oc <- fs_mr_oc_summary("fs_maxeffCons_mr_md40_knoise0_n500_res_1_1000.rds")
oc$estimation[oc$estimation$block == "H", c("estimator", "avg", "bias_beta")]
oc$identification
## End(Not run)
Enumerate the candidate family on a DGM's population frame
Description
Builds the family of candidate subgroups that forestsearch
would search over, but on the DGM's super-population frame
dgm$df_super rather than on a sampled trial, so that every cut lands
at a population quantile and every prevalence, purity and overlap is a
population proportion. The result carries the nine quantities the
operating-characteristics prediction of fs_oc_predict consumes.
Usage
fs_oc_family_enumerate(
dgm,
forestsearch_args,
n,
max_M = 2000L,
verbose = FALSE
)
Arguments
dgm |
An object of class |
forestsearch_args |
Named list of |
n |
Integer. Trial size at which the family is evaluated: sets the
size floor |
max_M |
Integer. Size guard: if more than |
verbose |
Logical. Print the enumeration counts at each stage. |
Details
Cuts. The cut-related entries of forestsearch_args
(confounders.name, conf.cont_jcuts, cut_type,
cont.cutoff, conf.cont_medians, conf.cont_medians_force,
conf_force, defaultcut_names, exclude_cuts,
collapse_cuts, collapse_cuts_args) are handed to the package's
own get_FSdata() on df_super. Both directions of every cut are
generated, as the search does. LASSO, GRF and DINA cut sources are off: they
are sample-fitted screens, not part of the cut specification.
Combinations. All combinations of up to maxk indicator
columns are enumerated with generate_combination_indices(), in the
order forestsearch() uses when it composes the MR family.
Structural floors (all on the population frame, in the order of
evaluate_combination_with_status()):
combinations with a constant column or an empty pairwise intersection are skipped (the search's status 0);
-
minp: every constituent factor must have population prevalence>= minp; -
rmin: each added factor must shrink the membership by more thanrminsubjects of a trial of size n, i.e. by more thanrmin / nin population proportion; size:
Pg >= n.min / n, wheren.minis resolved asforestsearch()resolves it – the supplied value (default 60), ormax(60, ceiling(n.min.frac * n))whenn.min = NULL.
The GRF variable-importance pre-screen, the effect screen, the consistency screen and the near-duplicate removal are not applied. Candidates whose population membership vectors are identical are collapsed to the first one enumerated.
Null DGMs. A DGM whose df_super$flag_harm has no member
(Q empty; generate_glm_dgm(model = "null")) is detected
structurally, cross-checked against dgm$model when present. Under
the null every candidate has the same true effect, so the fields become:
beta_g = the common effect (the DGM's effect_Qc =
effect_ITT, oriented as the alternative path orients);
se_g from the whole-population effective variance
(fs_dgm_scale(dgm, regions = list(S = ...)), the S row) at
(n, Pg) with the same prevalence scaling; PQg = 0;
sens_g = NA (0/0, undefined – not zero); spec_g = 1 - Pg;
PQ = 0 (from which NPV = 1 follows downstream). Enumeration, floors,
ovl and the covariance are unchanged. The element null
records the branch taken.
Scale. Every mean and standard error is derived from
fs_dgm_scale(dgm). The orientation is the harm direction
s = sign(m_tau[Q]), so the planted effect is oriented positive:
with Q the true harm region, tauQc = s * m_tau[Qc],
bint = s * (m_tau[Q] - m_tau[Qc]),
seQ1000 = sqrt(V_eff[Q] / (1000 * P(Q))), and for each candidate
beta_g = tauQc + bint * PQg – the signed mixture
s * (m_tau[Qc] + (m_tau[Q] - m_tau[Qc]) * PQg) – and
se_g = seQ1000 * sqrt(1000 / n) * sqrt(P(Q) / Pg) (sign-free, from
V_eff[Q]). Opposite-sign families (sign(m_tau[Qc]) != s)
are supported: benefit-direction candidates carry oriented-negative
beta_g, and tauQc may be negative for such families. When
both region effects share a sign the values coincide exactly with the
former oriented-absolute reading. A DGM with m_tau[Q] exactly zero
is rejected (no harm direction to orient by): plant a nonzero Q effect or
use the null path.
Value
An object of class c("fs_oc_family", "list") with elements
labCharacter, length M: the rule of each candidate.
PgPopulation prevalence
P(g).PQgPurity
P(g & Q) / P(g).beta_gOriented mixture mean
s * (m_tau[Qc] + (m_tau[Q] - m_tau[Qc]) * PQg), withs = sign(m_tau[Q])the harm direction.se_gAnchored standard error at
n.sens_gP(g & Q) / P(Q).spec_g1 - P(g & Qc) / P(Qc).ovlM-by-M matrix of
P(g_i & g_j).MNumber of candidates.
PQP(Q), the prevalence of the true region.membLogical matrix,
nrow(df_super)by M, the population membership of each candidate.nullLogical:
TRUEwhen the null branch was taken (Q empty).orientationAlternative branch only: list with the harm-direction sign
s, the signed region effectsm_tau_Qandm_tau_Qc, and the oriented mixture coefficientstauQc = s * m_tau_Qc(may be negative for opposite-sign families) andbint = s * (m_tau_Q - m_tau_Qc). Absent on the null branch.scaleThe
fs_dgm_scaleobject used.nThe trial size.
args_usedThe
forestsearch_argsentries consumed, with the resolvedn.min.cutsThe cut expressions
get_FSdata()produced.countsNamed integer vector: candidates at each stage.
See Also
Examples
set.seed(1)
N <- 2000
age <- round(rnorm(N, 35, 9)); pre <- round(rexp(N, 1 / 500))
V <- factor(rbinom(N, 1, 0.42), levels = 0:1)
inQ <- as.integer(age > 34 & pre <= 745)
mu0 <- 40 + 0.2 * age
dgm <- structure(list(
df_super = data.frame(age = age, preanti = pre, V = V,
mu0 = mu0, mu1 = mu0 - 26 - 14 * inQ,
flag_harm = inQ),
outcome_type = "continuous", effect_measure = "MD",
model_params = list(sigma = 127.5)), class = c("glm_dgm", "list"))
fam <- fs_oc_family_enumerate(
dgm, list(confounders.name = c("age", "preanti", "V"),
conf.cont_jcuts = list(age = 4, preanti = 4), n.min = 60),
n = 500)
fam
Sweep the declaration thresholds on one draw set per trial size
Description
Evaluates every fs_oc_predict() quantity over the full crossing of
n, c1 and c2, enumerating the family and drawing the
candidates' fitted effects once per n (and per gate) and
sweeping the thresholds against those draws. Thresholds enter only the
gate, so the sweep costs arithmetic, not draws.
Usage
fs_oc_grid(
dgm = NULL,
forestsearch_args = list(),
n,
c1,
c2,
family = NULL,
consistency_method = c("resample", "split"),
pconsistency = NULL,
draws = 2e+05,
block = 50000,
seed = NULL,
verbose = FALSE,
...
)
Arguments
dgm |
An object of class |
forestsearch_args |
Named list of |
n |
Numeric vector of trial sizes. |
c1, c2 |
Numeric vectors of screening and consistency floors; the grid
is their full crossing with |
family |
|
consistency_method |
|
pconsistency |
Numeric in (0, 1). Consistency-rate threshold
for the |
draws |
Integer. Number of Monte-Carlo draws. |
block |
Integer or |
seed |
Integer or |
verbose |
Logical. Report per- |
... |
Passed to |
Details
For each n the family is enumerated with
fs_oc_family_enumerate (or the supplied family is used
when its n matches), the draw set is generated under seed
(the same seed for every n: common random numbers across trial
sizes), and every (c1, c2) is evaluated on it. With
block = Inf the whole draw set is held and the computation is exactly
fs_oc_predict's – one grid point is identical() to a
call of fs_oc_predict() at the same settings and seed. With a finite
block the draws are generated in row blocks of that size and the
quantities accumulated, so memory is O(block x M) instead of O(draws x M);
blocked results agree with the one-block results to Monte-Carlo precision
(the RNG stream is laid out differently and proportions are formed as
count / draws), not bit-for-bit.
Value
An object of class c("fs_oc_grid", "list"):
tableData frame, one row per
(n, gate, c1, c2):n, consistency_method, pconsistency, c1, c2, M, draws, block, seed, det_rate, det_rate_se, EnH, Esens, Espec, Eppv, Enpv, EbetaH, Enaive_bias, mass_below.resultsList, parallel to the rows of
table, of the full per-point objects (classfs_oc_predict, withP1,p_sel,sel_cand their MC SEs).familiesPer
n:M, the floor, the stage counts and the family.timingPer
nand gate: seconds for enumeration, drawing and sweeping.settingsThe call's settings.
See Also
Examples
piQ <- 0.34
fam <- structure(list(
lab = c("Q", "P", "D"), Pg = c(piQ, 0.45, 0.31),
PQg = c(1, 0.28 / 0.45, 1), sens_g = c(1, 0.28 / piQ, 0.31 / piQ),
spec_g = c(1, 1 - 0.17 / (1 - piQ), 1),
ovl = matrix(c(piQ, 0.28, 0.31, 0.28, 0.45, 0.28 * 0.31 / piQ,
0.31, 0.28 * 0.31 / piQ, 0.31), 3, 3),
M = 3L, PQ = piQ, n = 500), class = c("fs_oc_family", "list"))
fam$beta_g <- 26 + 14 * fam$PQg
fam$se_g <- 13.7 * sqrt(2) * sqrt(piQ / fam$Pg)
g <- fs_oc_grid(family = fam, n = 500, c1 = c(20, 30, 40), c2 = 10,
consistency_method = "resample", draws = 2e4, seed = 1)
g
Invert the family declaration rate for a target
Description
Finds the threshold – c1 at fixed c2, or c2 at fixed
c1 – at which the family declaration rate equals target, on
one fixed draw set: "the c1 giving 80\
read as a type-I error under a null DGM.
Usage
fs_oc_invert(
dgm = NULL,
forestsearch_args = list(),
n,
target,
solve_for = c("c1", "c2"),
c1 = NULL,
c2 = NULL,
family = NULL,
consistency_method = c("resample", "split"),
pconsistency = NULL,
draws = 2e+05,
seed = NULL,
tol = 0.001,
...
)
Arguments
dgm |
An object of class |
forestsearch_args |
Named list of |
n |
Single trial size. |
target |
Numeric in (0, 1); may be a vector when
|
solve_for |
|
c1, c2 |
The fixed threshold (the one not solved for; the solved-for
argument is ignored). Defaults from |
family |
|
consistency_method |
One gate. |
pconsistency |
Numeric in (0, 1). Consistency-rate threshold
for the |
draws |
Integer. Number of Monte-Carlo draws. |
seed |
Integer or |
tol |
Bisection tolerance on the threshold scale ( |
... |
Passed to |
Details
Solving for c1: an order statistic. At fixed
(n, gate, c2) neither the eligible set nor the winner depends on
c1: with eligible_g = (Bhat_g - c2 >= z_p * se_g)
("resample") or (W1_g >= c2) & (W2_g >= c2) ("split"),
each draw's winner is w = argmax_{eligible} Bhat_g and its statistic
T = Bhat_w (-Inf when nothing is eligible). Declaration at
c1 is exactly T >= c1, and w is the maxeffCons winner
whenever anything declares. So the reduction is computed once
(.fs_oc_reduce()), the declaration rate at any c1 is
mean(T >= c1), and the largest c1 attaining a target is the
k-th largest T with k = ceiling(target * draws): an
order statistic, no search. The ceiling is mean(T > -Inf),
the fraction of draws with any eligible member, set by the consistency
screen alone; a target above it is unattainable and returns
NA with the ceiling and the binding threshold named – no
extrapolation. target may be a vector: one reduction serves every
target.
Solving for c2 at fixed c1: the eligible set moves
with c2, so the rate is bracketed from the draws' range and bisected
to tol on fixed draws (a monotone non-increasing step function); the
returned value is the largest threshold whose rate is still at least
target.
Every result reports the achieved rate and its MC SE, so the resolution the draw count supports is visible.
Value
For a single target, an object of class
c("fs_oc_invert", "list"): value (the threshold, or
NA), achieved (rate at value), achieved_se,
target, ceiling, ceiling_se, binding (which
threshold sets the ceiling), attainable, solve_for,
fixed, k (the order-statistic rank, solve_for = "c1"),
next_step_rate, iterations (bisections, solve_for =
"c2"), settings. For a vector target, a list of such
objects (class fs_oc_invert_list) with a table attribute.
See Also
Examples
piQ <- 0.34
fam <- structure(list(
lab = c("Q", "P", "D"), Pg = c(piQ, 0.45, 0.31),
PQg = c(1, 0.28 / 0.45, 1), sens_g = c(1, 0.28 / piQ, 0.31 / piQ),
spec_g = c(1, 1 - 0.17 / (1 - piQ), 1),
ovl = matrix(c(piQ, 0.28, 0.31, 0.28, 0.45, 0.28 * 0.31 / piQ,
0.31, 0.28 * 0.31 / piQ, 0.31), 3, 3),
M = 3L, PQ = piQ, n = 500), class = c("fs_oc_family", "list"))
fam$beta_g <- 26 + 14 * fam$PQg
fam$se_g <- 13.7 * sqrt(2) * sqrt(piQ / fam$Pg)
fs_oc_invert(family = fam, n = 500, target = 0.5, solve_for = "c1", c2 = 10,
consistency_method = "resample", draws = 2e4, seed = 1)
Predicted operating characteristics of the search over a candidate family
Description
Monte-Carlo prediction of the maxeffCons search's operating characteristics
– family declaration rate, per-candidate declaration, the selection
distribution, expected selected-subgroup size, sensitivity, specificity,
PPV, NPV, the oriented effect on the selected rule and the naive selection
bias – from the joint normal law of the candidates' fitted effects. The
family is a population enumeration from fs_oc_family_enumerate
or a supplied fs_oc_family object.
Usage
fs_oc_predict(
dgm = NULL,
forestsearch_args = list(),
n,
c1 = NULL,
c2 = NULL,
family = NULL,
consistency_method = c("resample", "split"),
pconsistency = NULL,
draws = 2e+05,
seed = NULL,
...
)
Arguments
dgm |
An object of class |
forestsearch_args |
Named list of |
n |
Integer. Trial size. Overrides any size implied by the arguments; sets the size floor and the standard-error scale of an enumerated family and converts expected prevalence to expected subjects. |
c1 |
Numeric. Screening floor on the full-sample effect – the |
c2 |
Numeric. Consistency floor on each half-sample effect – the
|
family |
|
consistency_method |
|
pconsistency |
Numeric in (0, 1). Consistency-rate threshold
for the |
draws |
Integer. Number of Monte-Carlo draws. |
seed |
Integer or |
... |
Passed to |
Details
The full-sample fitted effects of the M candidates are taken as
N(\beta_g, S_g) with S_g = \rho \circ (se_g se_g'),
\rho_{ij} = P(g_i \cap g_j) / \sqrt{P(g_i) P(g_j)}. Two independent
half-sample draws W1, W2, each with covariance 2 * Sg,
are generated through fs_sym_root; the full-sample effect is
Bhat = (W1 + W2) / 2.
The two gates. Family declaration is any candidate declaring; the selected rule is the effect maximiser among declaring candidates (maxeffCons).
"resample"– the package's production screenThe consistency stage of
forestsearch(defaultconsistency_method = "resample") represents the random 50/50 split's half effects as\hat\beta \pm Dand computes the consistency rate in closed form as2\Phi((\hat\beta - c_2)/\sigma_D) - 1, a candidate passing when that rate is at leastpconsistency(R/consistency_resample.R;R/subgroup_consistency_helpers.Rdrops a candidate whenp.consistency < pconsistency.threshold). Inverting:\hat\beta \ge c_2 + z_p \sigma_Dwithz_p = \Phi^{-1}((1+p)/2).\sigma_Dissqrt(sum(dfbeta[i, treat]^2)), the sandwich standard error of the subgroup treatment coefficient; in the draw model above the analogous quantity isse_g, and on a simulated MD40 trial the two agree within a few percent with no prevalence trend, so the wrapper identifies\sigma_D = se_g. The gate is then a single threshold on the full-sample draw,(Bhat >= c1) & (Bhat - c2 >= z_p * se_g), and onlyBhat ~ N(beta_g, Sg)is drawn – one matrix, not two."split"– the analytic document's historical gateOne half-sample pair
W1,W2stands in for the consistency rate:(W1 + W2 >= 2 * c1) & (W1 >= c2) & (W2 >= c2), the screening floorc1on the full-sample effect and the consistency floorc2on each half. Kept bit-identical to the document'sworked-predictionschunk.
All expectations are selection-weighted population functionals of the
family, conditional on declaration. Monte-Carlo standard errors are given
for the proportions (sqrt(p (1 - p) / draws)).
Value
An object of class c("fs_oc_predict", "list"):
det_rate,det_rate_seFamily declaration rate and its MC standard error.
P1,P1_sePer-candidate declaration probability.
p_sel,p_sel_sePer-candidate selection probability (unconditional).
sel_cSelection distribution given declaration.
EnHExpected selected-subgroup size in subjects,
n * sum(sel_c * Pg).Esens,Espec,Eppv,EnpvExpected classification metrics of the selected rule against Q.
EbetaHExpected oriented true effect on the selected rule.
Enaive_biasExpected naive minus true effect on the selected rule.
mass_belowSelection mass on rules whose true mean is below
c1.M,labFamily size and labels.
settingsn,c1,c2,consistency_method,pconsistency(NAfor"split"),draws.seedThe seed used.
familyThe family object.
See Also
fs_oc_family_enumerate, fs_sym_root
Examples
# A hand-built three-candidate family
piQ <- 0.34
fam <- structure(list(
lab = c("Q", "P", "D"), Pg = c(piQ, 0.45, 0.31),
PQg = c(1, 0.28 / 0.45, 1), sens_g = c(1, 0.28 / piQ, 0.31 / piQ),
spec_g = c(1, 1 - 0.17 / (1 - piQ), 1),
ovl = matrix(c(piQ, 0.28, 0.31, 0.28, 0.45, 0.28 * 0.31 / piQ,
0.31, 0.28 * 0.31 / piQ, 0.31), 3, 3),
M = 3L, PQ = piQ), class = c("fs_oc_family", "list"))
fam$beta_g <- 26 + 14 * fam$PQg
fam$se_g <- 13.7 * sqrt(2) * sqrt(piQ / fam$Pg)
fs_oc_predict(family = fam, n = 500, c1 = 30, c2 = 10,
consistency_method = "resample", draws = 2e4, seed = 1)
fs_oc_predict(family = fam, n = 500, c1 = 30, c2 = 10,
consistency_method = "split", draws = 2e4, seed = 1)
Bias-vs-coverage display (three panels)
Description
Draws the standard bias-vs-coverage display from a stacked
fs_sim_bias_coverage() table carrying a cell column: (1) one-sided
coverage against b with Gaussian-reference curves Phi(z1 * r - b) for
the given curves values of r; (2) two-sided coverage against b with
curves Phi(z2 * r - b) - Phi(-z2 * r - b); (3) observed coverage against
the Gaussian reference at each point's own (b, r), with the diagonal.
Estimators are distinguished by marker shape, the nominal level by a
dashed line, and cells by text labels.
Usage
fs_plot_bias_coverage(
tbl,
labels = TRUE,
curves = c(1, 1.25, 1.5, 2),
level = 0.95,
side = c("lower", "upper")
)
Arguments
tbl |
Stacked output of |
labels |
Draw cell labels next to the points (default |
curves |
Gaussian-reference |
level |
Nominal level for the reference lines (default 0.95). |
side |
|
Value
A patchwork object of the three ggplot panels.
Covariate signatures of the realized rules
Description
Collapses realized rules to the set of covariates they name, discarding
the cut values. A rules-by-frequency table lists every distinct sg_def
string and so is dominated by cut noise – a 500-replicate survival run
produced 583 rows, in which {age <= 46} & {nodes <= 5} and
{nodes <= 3} & {age <= 46} are separate entries despite being the same
pairing. This reports how often each covariate combination is selected,
which is the structural question.
Usage
fs_rule_covariate_pairs(x, true_covariates = NULL, detected_only = TRUE)
Arguments
x |
A results data frame carrying |
true_covariates |
Character vector. The true-region covariates, used to
classify each signature as |
detected_only |
Logical. Restrict to rows with |
Value
A data frame ordered by frequency: covariates, k, n, share,
n_rules, and match when true_covariates is given.
What a signature is
The covariate names in the rule, de-duplicated and sorted, joined by " & ".
Sorting is what makes the two orderings above collapse together. Two
consequences worth knowing:
A rule naming one covariate twice — a band such as
!{nodes <= 3} & {nodes <= 5}— has signature"nodes", of size 1. Thekcolumn gives the number of distinct covariates, andn_rulesthe number of distinct rule strings that collapsed into the row, so bands are visible rather than hidden.Logically equivalent spellings of the same condition (
!{meno}versus{meno == 0}) collapse together, which a raw rule table also splits.
Estimator Standard Deviation at a Given Trial Size
Description
Converts a fs_dgm_scale object into the sampling standard
deviation of the within-region difference-in-means estimator at trial size
n.
Usage
fs_scale_se(scale, n, region = NULL, jensen = FALSE)
Arguments
scale |
An object of class |
n |
Integer. Trial size. |
region |
Character. Region name, matching a row of
|
jensen |
Logical. If |
Value
Named numeric vector of standard deviations.
See Also
Bias and coverage summary for one simulation cell
Description
From a combined bundle's per-replicate results (the recorder columns of
the fb_mr_field template family: nv_H_*, mr_H_*, fld_H_*,
betaHhat_H, detected, and their Hc twins), computes, per estimator
over the detected replicates: the retained bias on the log scale, the
empirical SD of the log estimate, the mean reported SE, their ratios
b = bias_log / sd_emp and r = se_mean / sd_emp, the observed one- and
two-sided coverage of the target with Wilson intervals, and the Gaussian
reference coverages at (b, r).
Usage
fs_sim_bias_coverage(
results,
block = c("H", "Hc"),
estimators = c("naive", "mr", "fld"),
level = 0.95,
target = "betaHhat",
side = c("lower", "upper"),
scale = c("log", "identity")
)
Arguments
results |
Data frame of per-replicate records from a combined bundle
( |
block |
|
estimators |
Subset of |
level |
Coverage level (default 0.95). |
target |
|
side |
|
scale |
|
Details
Conventions: the
one-sided lower bound is the stored fld_<block>_lo1s for the field row
and exp(log est - qnorm(level) * SE) for naive and MR (IJ); two-sided
coverage uses each estimator's stored bounds; the field estimator is the
shrunk-field est2 with the two-sided Lambda-quantile interval. Under
side = "upper" (the complement's benefit-claim orientation) the one-sided coverage is
truth <= upper bound, with the field's stored fld_<block>_up1s and
exp(log est + qnorm(level) * SE) for naive and MR (IJ), and the Gaussian
reference is Phi(z_level * r + b) – a positive retained bias helps an
upper bound where it hurts a lower one.
Value
Data frame, one row per estimator: estimator, n, bias_log,
sd_emp, se_mean, b, r, cov1, cov1_wilson_lo,
cov1_wilson_hi, cov2, cov2_wilson_lo, cov2_wilson_hi,
cov1_ref, cov2_ref.
Symmetric Square Root of a Covariance Matrix
Description
Computes the symmetric matrix square root of scale * S after
symmetrising S as (S + t(S)) / 2: with eigendecomposition
S = V D V', the returned matrix is R = V D^{1/2} V' (eigenvalues
scaled by scale and clamped at zero before the square root), so that
R R' = \mathrm{scale} \cdot S.
Usage
fs_sym_root(S, scale = 2)
Arguments
S |
Numeric square matrix, a covariance (or any matrix whose
symmetrisation |
scale |
Numeric scalar multiplying |
Value
A symmetric numeric matrix R of the same dimension as
S, satisfying R %*% t(R) equal to scale times the
symmetrised S (up to the zero-clamp on negative eigenvalues).
Why the symmetric root and not V D^(1/2)
Both roots satisfy R R' = \mathrm{scale} \cdot S, but the asymmetric
root V D^{1/2} depends on the eigenvector basis, which is not a
continuous function of its input: signs flip and vectors rotate freely
inside degenerate subspaces – which a candidate covariance built from
complement pairs has exactly, since the complement constraints make it
rank-deficient by construction. A pinned RNG seed fixes the standard-normal
draws Z; it cannot fix the basis Z is multiplied by, so the asymmetric root
is non-reproducible across arithmetically equivalent routes to the same
matrix. The symmetric root V D^{1/2} V' is a matrix function of the
symmetrised input alone – basis-free and continuous – so samples built
from it are continuous, reproducible functions of the inputs.
Examples
S <- crossprod(matrix(rnorm(25), 5)) # symmetric positive definite
R <- fs_sym_root(S, scale = 2)
max(abs(R %*% t(R) - 2 * S)) # ~ 1e-14
isSymmetric(R)
# Rank-deficient input: the root exists and stays symmetric
X <- matrix(rnorm(15), 5, 3)
R2 <- fs_sym_root(tcrossprod(X), scale = 1)
max(abs(R2 %*% t(R2) - tcrossprod(X)))
Run Guo and He Algorithm 3 on a forestsearch maxeff fit
Description
Bridge from a forestsearch(sg_focus = "maxeff") fit to the Guo and He
(2021) Algorithm 3 de-biasing estimator guohe_algorithm3(). This is the
"post-hoc identified subgroups" application of Guo and He: forestsearch
enumerates a large candidate family and selects the argmax; Guo and He
supplies a de-biased effect and one-sided bound for the selected subgroup,
correcting the winner's-curse optimism of that selection. The candidate
family is materialized from the fit with fs_export_maxeff_family() and,
by default, verified against forestsearch's own selected-subgroup
membership with fs_assert_membership() before any de-biasing runs.
Usage
fs_to_guohe(
fit,
time = "time_months",
event = "status",
treatment = "hormon",
B = 2000L,
r = 0.03,
level = 0.05,
seed = NULL,
min_events = 5L,
parallel = FALSE,
verify = TRUE
)
Arguments
fit |
A |
time, event, treatment |
Names of the time, event, and 0/1 treatment
columns on |
B |
Number of bootstrap resamples passed to |
r |
Shrinkage tuning parameter, strictly between 0 and 0.5; see
|
level |
One-sided level; the two-sided interval has coverage
|
seed |
Optional integer for reproducible resampling. |
min_events |
Minimum events for a candidate to be estimable. |
parallel |
Logical; passed to |
verify |
Logical; if |
Details
The de-biasing engine is an independent implementation of the published method, validated by reproducing the simulation study of Section 5 of Guo and He (2021).
The Guo and He selection is orient = +1 (the maximum-hazard-ratio, i.e.
harm, subgroup), so its argmax coincides with the forestsearch maxeff
selection. If the two nevertheless disagree (possible when candidates are
dropped as non-estimable), a warning is issued and the returned Guo and He
objects refer to the Guo and He argmax.
Value
A list with elements gh (the full "guohe_a3" object), export
(the fs_export_maxeff_family() result), selected (the Guo and He
argmax candidate name), and the convenience scalars naive_hr,
debiased_hr, and bound_hr (hazard-ratio scale).
References
Guo, X. and He, X. (2021). Inference on Selected Subgroups in Clinical Trials. Journal of the American Statistical Association, 116(535), 1498–1506. doi:10.1080/01621459.2020.1740096
See Also
fs_export_maxeff_family(), fs_assert_membership(),
guohe_algorithm3(), guohe_adaptive_r()
Examples
set.seed(11)
n <- 200
df <- data.frame(id = seq_len(n), treat = rbinom(n, 1, 0.5),
age = round(rnorm(n, 55, 12)),
bm = factor(rbinom(n, 1, 0.5)))
harm <- df$age <= 50 & df$bm == "1"
tt <- rexp(n, 0.05 * exp(log(2.5) * df$treat * harm))
df$time <- pmin(tt, 60)
df$event <- as.integer(tt <= 60)
fit <- forestsearch(
df.analysis = df,
outcome.name = "time", event.name = "event",
treat.name = "treat", id.name = "id",
confounders.name = c("age", "bm"),
sg_focus = "maxeff", use_grf = FALSE, use_lasso = FALSE,
fs.splits = 50, n.min = 30, d0.min = 5, d1.min = 5,
hr.threshold = 1.1, hr.consistency = 1.0,
pconsistency.threshold = 0.5, maxk = 2,
details = FALSE, plot.sg = FALSE,
parallel_args = list(plan = "sequential", workers = 1L,
show_message = FALSE)
)
gh <- fs_to_guohe(fit, time = "time", event = "event",
treatment = "treat", B = 50, seed = 3)
gh$naive_hr
gh$debiased_hr
gh$bound_hr
Generate Synthetic Survival Data using AFT Model with Flexible Subgroups
Description
Creates a data generating mechanism (DGM) for survival data using an Accelerated Failure Time (AFT) model with Weibull distribution. Supports flexible subgroup definitions and treatment-subgroup interactions.
Usage
generate_aft_dgm_flex(
data,
continuous_vars,
factor_vars,
continuous_vars_cens = NULL,
factor_vars_cens = NULL,
set_beta_spec = list(set_var = NULL, beta_var = NULL),
outcome_var,
event_var,
treatment_var = NULL,
subgroup_vars = NULL,
subgroup_cuts = NULL,
draw_treatment = FALSE,
model = "alt",
k_treat = 1,
k_inter = 1,
n_super = 5000,
select_censoring = TRUE,
cens_type = "weibull",
cens_params = list(),
cens_intercept_only = FALSE,
seed = 8316951,
verbose = TRUE,
standardize = FALSE,
spline_spec = NULL
)
Arguments
data |
A data.frame containing the input dataset to base the simulation on |
continuous_vars |
Character vector of continuous variable names to be standardized and included as covariates |
factor_vars |
Character vector of factor/categorical variable names to be converted to dummy variables (largest value as reference) |
continuous_vars_cens |
Character vector of continuous variable names to be used for censoring model. If NULL, uses same as continuous_vars. Default NULL |
factor_vars_cens |
Character vector of factor variable names to be used for censoring model. If NULL, uses same as factor_vars. Default NULL |
set_beta_spec |
List with elements 'set_var' and 'beta_var' for manually setting specific beta coefficients. Default list(set_var = NULL, beta_var = NULL) |
outcome_var |
Character string specifying the name of the outcome/time variable |
event_var |
Character string specifying the name of the event/status variable (1 = event, 0 = censored) |
treatment_var |
Character string specifying the name of the treatment variable. If NULL, treatment will be randomly simulated with 50/50 allocation |
subgroup_vars |
Character vector of variable names defining the subgroup. Default is NULL (no subgroups) |
subgroup_cuts |
Named list of cutpoint specifications for subgroup variables. See Details section for flexible specification options |
draw_treatment |
Logical indicating whether to redraw treatment assignment in simulation. Default is FALSE (use original assignments) |
model |
Character string: "alt" for alternative model with subgroup effects, "null" for null model without subgroup effects. Default is "alt" |
k_treat |
Numeric treatment effect modifier. Values >1 increase treatment effect, <1 decrease it. Default is 1 (no modification) |
k_inter |
Numeric interaction effect modifier for treatment-subgroup interaction. Default is 1 (no modification) |
n_super |
Integer specifying size of super population to generate. Default is 5000 |
select_censoring |
Logical. If |
cens_type |
Character string specifying censoring distribution type:
|
cens_params |
Named list of censoring distribution parameters.
Interpretation depends on
Default |
cens_intercept_only |
Logical. Honored only in force-fit mode
(
Setting |
seed |
Integer random seed for reproducibility. Default is 8316951 |
verbose |
Logical indicating whether to print diagnostic information during execution. Default is TRUE |
standardize |
Logical indicating whether to standardize continuous variables. Default is FALSE |
spline_spec |
List specifying spline configuration for treatment effect. Must include 'var' (variable name), 'knot', 'zeta', and 'log_hrs' (vector of length 3). Default NULL (no spline) |
Details
Subgroup Cutpoint Specifications
The subgroup_cuts parameter accepts multiple flexible specifications:
Fixed Value
subgroup_cuts = list(er = 20) # er <= 20
Quantile-based
subgroup_cuts = list( er = list(type = "quantile", value = 0.25) # er <= 25th percentile )
Function-based
subgroup_cuts = list( er = list(type = "function", fun = median) # er <= median )
Range
subgroup_cuts = list( age = list(type = "range", min = 40, max = 60) # 40 <= age <= 60 )
Greater than
subgroup_cuts = list( nodes = list(type = "greater", quantile = 0.75) # nodes > 75th percentile )
Multiple values (for categorical)
subgroup_cuts = list( grade = list(type = "multiple", values = c(2, 3)) # grade in (2, 3) )
Custom function
subgroup_cuts = list(
er = list(
type = "custom",
fun = function(x) x <= quantile(x, 0.3) | x >= quantile(x, 0.9)
)
)
Model Structure
The AFT model with Weibull distribution is specified as:
\log(T) = \mu + \gamma' X + \sigma \epsilon
Where:
-
Tis the survival time -
\muis the intercept -
\gammacontains the covariate effects -
Xincludes treatment, covariates, and treatment x subgroup interaction -
\sigmais the scale parameter -
\epsilonfollows an extreme value distribution
Interaction Term
The model creates a SINGLE interaction term representing the treatment effect modification when ALL subgroup conditions are simultaneously satisfied. This is not multiple separate interactions but one combined indicator.
Value
An object of class "aft_dgm_flex" (a list) with components:
df_superData frame containing the super-population, including covariates, treatment, counterfactual linear predictors, and subgroup indicator
flag_harm.df_sourceData frame with every source patient exactly once (the prepared trial data), carrying the same counterfactual outcome and censoring linear predictors as
df_super. Consumed bysimulate_from_dgm(baseline = "fixed")to hold baseline covariates – and hence subgroup composition – fixed across simulated trials.model_paramsList with AFT parameters
mu,tau,gamma,b0, the fittedcensoringmodel, and optionalspline_info.subgroup_infoList describing the true subgroup:
vars,cuts,definitions,size,proportion.hazard_ratiosList of true HR/AHR/CDE values on the super-population (see
compute_dgm_cde).analysis_varsNamed list of column roles (continuous, factor, covariates, treatment, outcome, event).
model_type"null"or"alt".n_superSuper-population size.
seedSeed used for super-population generation.
Author(s)
Your Name
References
Leon, L.F., et al. (2024). Statistics in Medicine.
Kalbfleisch, J.D. and Prentice, R.L. (2002). The Statistical Analysis of Failure Time Data (2nd ed.). Wiley.
Examples
## Not run:
df <- survival::gbsg
dgm <- generate_aft_dgm_flex(
data = df,
outcome_var = "rfstime",
event_var = "status",
treatment_var = "hormon",
continuous_vars = c("age", "size", "nodes", "pgr", "er"),
factor_vars = "meno",
model = "null",
verbose = FALSE
)
str(dgm)
## End(Not run)
Generate Synthetic Data using Bootstrap with Perturbation
Description
Generate Synthetic Data using Bootstrap with Perturbation
Usage
generate_bootstrap_synthetic(
data,
continuous_vars,
cat_vars,
n = NULL,
seed = 123,
noise_level = 0.1,
id_var = NULL,
cat_flip_prob = NULL,
preserve_bounds = TRUE,
ordinal_vars = NULL
)
Arguments
data |
Original dataset to bootstrap from |
continuous_vars |
Character vector of continuous variable names |
cat_vars |
Character vector of categorical variable names |
n |
Number of synthetic observations to generate (default: same as original) |
seed |
Random seed for reproducibility |
noise_level |
Noise level for perturbation (0 to 1, default 0.1) |
id_var |
Optional name of ID variable to regenerate (will be numbered 1:n) |
cat_flip_prob |
Probability of flipping categorical values (default: noise_level/2) |
preserve_bounds |
Logical: should continuous variables stay within original bounds? (default: TRUE) |
ordinal_vars |
Optional character vector of ordinal categorical variables (these will be perturbed to adjacent values rather than randomly flipped) |
Value
A data frame with synthetic data
Examples
# Example 1: Using with GBSG dataset
synth_gbsg <- generate_bootstrap_synthetic(
data = survival::gbsg,
continuous_vars = c("age", "size", "nodes", "pgr", "er", "rfstime"),
cat_vars = c("meno", "hormon", "status"),
ordinal_vars = c("grade"),
id_var = "pid",
n = 1000,
seed = 123,
noise_level = 0.15
)
# Example 2: Using with any dataset
my_data <- data.frame(
id = 1:100,
height = rnorm(100, 170, 10),
weight = rnorm(100, 70, 15),
age = sample(20:80, 100, replace = TRUE),
gender = sample(c("M", "F"), 100, replace = TRUE),
education = sample(1:5, 100, replace = TRUE),
smoker = sample(0:1, 100, replace = TRUE)
)
synth_data <- generate_bootstrap_synthetic(
data = my_data,
continuous_vars = c("height", "weight", "age"),
cat_vars = c("gender", "smoker"),
ordinal_vars = c("education"),
id_var = "id",
n = 150,
seed = 456
)
Generate Bootstrap Sample with Added Noise
Description
Creates a bootstrap sample from a dataset with controlled noise added to both continuous and categorical variables. This function is useful for generating synthetic datasets that maintain the general structure of the original data while introducing controlled variation.
Usage
generate_bootstrap_with_noise(
data,
n = NULL,
continuous_vars = NULL,
cat_vars = NULL,
id_var = "pid",
seed = 123,
noise_level = 0.1
)
Arguments
data |
A data frame containing the original dataset to bootstrap from. |
n |
Integer. Number of observations in the output dataset. If NULL (default), uses the same number of rows as the input data. |
continuous_vars |
Character vector of column names to treat as continuous variables. If NULL (default), automatically detects numeric columns. |
cat_vars |
Character vector of column names to treat as categorical variables. If NULL (default), automatically detects factors, logical columns, and numeric columns with 10 or fewer unique values. |
id_var |
Character string specifying the name of the ID variable column. This column will be reset to sequential values (1:n) in the output. Default is "pid". |
seed |
Integer. Random seed for reproducibility. Default is 123. |
noise_level |
Numeric between 0 and 1. Controls the amount of noise added. For continuous variables, this is multiplied by the standard deviation to determine noise magnitude. For categorical variables, this is divided by 2 to determine the probability of value changes. Default is 0.1. |
Details
The function performs the following operations:
Bootstrap Sampling
Samples n observations with replacement from the original dataset.
Continuous Variable Noise
Adds Gaussian noise with standard deviation = original SD × noise_level
Constrains values to remain within original variable bounds
Preserves integer type for variables that appear to be integers
Categorical Variable Perturbation
Changes values with probability = noise_level / 2
Binary variables: flips to opposite value
Multi-level unordered: randomly selects from other levels
Ordered factors: weights selection toward adjacent levels
Preserves factor levels and ordering from original data
Value
A data frame with the same structure as the input data, containing bootstrap sampled observations with added noise.
Note
The function assumes that categorical variables with numeric encoding should maintain their numeric type unless they are factors in the input
Missing values (NA) are handled appropriately in calculations but are not imputed
For ordered factors or variables named "grade", the perturbation favors transitions to adjacent levels over distant levels
See Also
sample for bootstrap sampling,
rnorm for noise generation
Examples
## Not run:
# Load example dataset
data(gbsg, package = "survival")
# Basic usage with automatic variable detection
synthetic_data <- generate_bootstrap_with_noise(
data = gbsg,
seed = 123
)
# Specify variables explicitly
synthetic_data <- generate_bootstrap_with_noise(
data = gbsg,
n = 1000,
continuous_vars = c("age", "size", "nodes", "pgr", "er", "rfstime"),
cat_vars = c("meno", "grade", "hormon", "status"),
id_var = "pid",
seed = 456,
noise_level = 0.15
)
# Create multiple synthetic datasets
synthetic_list <- lapply(1:10, function(i) {
generate_bootstrap_with_noise(data = gbsg, seed = i)
})
## End(Not run)
Generate Combination Indices
Description
Creates indices for all factor combinations up to maxk
Usage
generate_combination_indices(L, maxk)
Generate Complement Expression
Description
Creates the logical complement of a subgroup expression. Handles common patterns like "var <= x" -> "var > x".
Usage
generate_complement_expression(expr)
Arguments
expr |
Character vector of expressions to negate. |
Value
Character string with negated expression.
Generate Detection Probability Curve
Description
Computes detection probability across a range of hazard ratios to create a power-like curve for subgroup detection.
Usage
generate_detection_curve(
theta_range = c(0.5, 3),
n_points = 50L,
n_sg,
prop_cens = 0.3,
hr_threshold = 1.25,
hr_consistency = 1,
include_reference = TRUE,
method = "cubature",
verbose = TRUE
)
Arguments
theta_range |
Numeric vector of length 2. Range of HR values to evaluate. Default: c(0.5, 3.0) |
n_points |
Integer. Number of points to evaluate. Default: 50 |
n_sg |
Integer. Subgroup sample size. |
prop_cens |
Numeric. Proportion censored (0-1). Default: 0.3 |
hr_threshold |
Numeric. HR threshold for detection. Default: 1.25 |
hr_consistency |
Numeric. HR consistency threshold. Default: 1.0 |
include_reference |
Logical. Include reference HR values (0.5, 0.75, 1.0). Default: TRUE |
method |
Character. Integration method. Default: "cubature" |
verbose |
Logical. Print progress. Default: TRUE |
Value
A data.frame with columns:
theta |
Hazard ratio values |
probability |
Detection probability |
n_sg |
Subgroup size (repeated) |
prop_cens |
Censoring proportion (repeated) |
hr_threshold |
Detection threshold (repeated) |
Examples
## Not run:
# Generate detection curve
curve_data <- generate_detection_curve(
n_sg = 60,
prop_cens = 0.2,
hr_threshold = 1.25
)
# Plot
plot_detection_curve(curve_data)
## End(Not run)
Generate Synthetic GBSG Data using Generalized Bootstrap
Description
Generate Synthetic GBSG Data using Generalized Bootstrap
Usage
generate_gbsg_bootstrap_general(n = 686, seed = 123, noise_level = 0.1)
Arguments
n |
Number of observations |
seed |
Random seed |
noise_level |
Noise level for perturbation |
Value
Synthetic GBSG dataset
Examples
## Not run:
df_synth <- generate_gbsg_bootstrap_general(n = 200, seed = 42)
dim(df_synth)
names(df_synth)
## End(Not run)
Generate GLM-Based Data Generating Mechanism
Description
Creates a data generating mechanism (DGM) for non-survival outcomes using generalised linear models.
Usage
generate_glm_dgm(
data,
factor_vars,
continuous_vars = NULL,
outcome_var,
treatment_var,
outcome_type = c("binary", "continuous", "count"),
effect_measure = NULL,
offset_var = NULL,
subgroup_vars = NULL,
subgroup_cuts = NULL,
model = c("alt", "null"),
k_treat = 1,
k_inter = 0,
adverse_outcome = FALSE,
n_super = 5000L,
seed = 8316951L,
verbose = FALSE,
k_random_noise = 0L,
noise_seed = 20260807L
)
Arguments
data |
A data frame containing the source dataset. |
factor_vars |
Character vector of factor/categorical variable names
to include as covariates. These are the candidate confounders for
ForestSearch (e.g., |
continuous_vars |
Character vector of continuous variable names
to include as prognostic covariates in the baseline GLM. These
improve the model fit but do not participate in subgroup definition.
Consistent with |
outcome_var |
Character string naming the outcome variable. |
treatment_var |
Character string naming the treatment variable (must be 0/1 coded). |
outcome_type |
Character. One of |
effect_measure |
Character. Effect measure for the GLM.
Default is |
offset_var |
Character or |
subgroup_vars |
Character vector of variable names defining the
subgroup. Default |
subgroup_cuts |
Named list of cutpoint specifications for
subgroup variables. Uses the same flexible format as
|
model |
Character. |
k_treat |
Numeric. Scaling factor for the fitted treatment
coefficient on the linear predictor scale. |
k_inter |
Numeric. Direct additive shift on the linear predictor
for Q members under treatment. The interpretation depends on
Default |
adverse_outcome |
Logical. If |
n_super |
Integer. Size of the super-population. Default
|
seed |
Integer. Random seed for super-population sampling.
Default |
verbose |
Logical. Print diagnostic information. Default
|
k_random_noise |
Integer. Number of standard-normal noise columns
( |
noise_seed |
Integer. Seed for the population-level noise draw, used
only when |
Details
Fits a GLM to the source data, creates a super-population with
individual-level potential outcomes, and stores true subgroup effects.
The returned object can be passed to simulate_from_glm_dgm
for simulation studies, and is compatible with
run_simulation_analysis (which dispatches on class).
Noise columns requested via k_random_noise are inert
N(0,1) population attributes intended as candidate confounders; the
builder never puts them in the outcome model. They are drawn once onto
df_super after sampling, so simulated trials inherit them by
row resampling.
Value
An object of class c("glm_dgm", "list") with:
df_superSuper-population data frame with potential outcome columns:
p0,p1(binary probabilities), ormu0,mu1(conditional means for continuous / count outcomes), andflag_harm.df_sourceThe observed analysis frame in observed order (every source patient exactly once), carrying the same potential-outcome columns as
df_supercomputed by the identical prediction calls, plusflag_harmand the discretised covariates. This is the frozen panel consumed bysimulate_from_glm_dgm(baseline = "fixed").hazard_ratiosNamed list with
harm_subgroup,no_harm_subgroup, andoverall– effect estimates on the scale determined byeffect_measure. Field name retained for compatibility withget_dgm_hr()and reporting functions.outcome_typeCharacter:
"binary","continuous", or"count".effect_measureCharacter: the effect measure used.
model_paramsList with fitted coefficients, family,
k_inter, residual SD (continuous), and offset variable name (count).subgroup_infoList with subgroup definition, size, and proportion.
model_typeCharacter:
"alt"or"null".noise_namesCharacter vector of noise column names (
character(0)whenk_random_noise = 0).noise_scheme"population"when noise was drawn, otherwise"none".noise_seedThe seed used for the noise draw, or
NA_integer_when none was drawn.
See Also
simulate_from_glm_dgm,
calibrate_glm_interaction,
generate_aft_dgm_flex
Examples
## Not run:
library(speff2trial)
actg <- subset(ACTG175, arms %in% c(1, 3))
actg$treat <- 1L - as.integer(actg$arms == 1)
actg$y <- as.integer(actg$cd420 > actg$cd40)
actg$z1 <- as.factor(ifelse(actg$age > 34, 1L, 0L))
actg$z2 <- as.factor(ifelse(actg$preanti <= 745, 1L, 0L))
dgm <- generate_glm_dgm(
data = actg,
factor_vars = c("z1", "z2"),
outcome_var = "y",
treatment_var = "treat",
outcome_type = "binary",
subgroup_vars = c("z1", "z2"),
subgroup_cuts = list(z1 = 1L, z2 = 1L),
model = "alt",
k_inter = 1.0, # additive log-odds shift for Q under treatment
verbose = TRUE
)
## End(Not run)
Generate Readable Subgroup Labels from ForestSearch Object
Description
Extracts human-readable subgroup labels that are also valid R expressions
for use with plotKM.band_subgroups(). Attempts to extract the actual
subgroup definition (e.g., "er <= 0") rather than column references.
Usage
generate_readable_sg_labels(fs.est, verbose = FALSE)
Arguments
fs.est |
A forestsearch object. |
verbose |
Logical. Print diagnostic messages. |
Value
Character vector of length 2: c(harm_label, benefit_label)
Generate Super Population and Calculate Linear Predictors
Description
Generate Super Population and Calculate Linear Predictors
Usage
generate_super_population(
df_work,
n_super,
draw_treatment,
gamma,
b0,
mu,
tau,
verbose,
spline_info = NULL
)
Fit Cox Model for Subgroup
Description
Fits a Cox model for a subgroup and returns estimate and standard error.
Usage
get_Cox_sg(
df_sg,
cox.formula,
est.loghr = TRUE,
cox_initial = log(1),
treat.name = NULL
)
Arguments
df_sg |
Data frame for subgroup. |
cox.formula |
Cox model formula. |
est.loghr |
Logical. Is estimate on log(HR) scale? |
cox_initial |
Optional pre-fitted Cox model object to use instead of fitting a new model. Default NULL |
treat.name |
Character or |
Details
Function is utilized throughout codebase
Value
List with estimate and standard error.
Examples
## Not run:
library(survival)
df <- data.frame(
tte = gbsg$rfstime / 30.4375,
event = gbsg$status,
treat = gbsg$hormon
)
formula <- build_cox_formula("tte", "event", "treat")
result <- get_Cox_sg(df, cox.formula = formula)
exp(result$est_obs) # hazard ratio
## End(Not run)
ForestSearch Data Preparation and Feature Selection
Description
Prepares a dataset for ForestSearch, including options for LASSO-based dimension reduction, GRF cuts, forced cuts, and flexible cut strategies. Returns a list with the processed data, subgroup factor names, cut expressions, and LASSO selection results.
Usage
get_FSdata(
df.analysis,
use_lasso = FALSE,
use_grf = FALSE,
grf_cuts = NULL,
confounders.name,
cont.cutoff = 4,
conf_force = NULL,
conf.cont_medians = NULL,
conf.cont_medians_force = NULL,
conf.cont_jcuts = NULL,
dina_cuts = NULL,
collapse_cuts = TRUE,
collapse_cuts_args = list(),
replace_med_grf = TRUE,
defaultcut_names = NULL,
cut_type = "default",
exclude_cuts = NULL,
outcome.name = "tte",
event.name = "event",
details = TRUE,
outcome_type = "survival",
offset.name = NULL
)
Arguments
df.analysis |
Data frame containing the data. |
use_lasso |
Logical. Whether to use LASSO for dimension reduction. |
use_grf |
Logical. Whether to use GRF cuts. |
grf_cuts |
Character vector of GRF cut expressions. |
confounders.name |
Character vector of confounder variable names. |
cont.cutoff |
Integer. Cutoff for continuous variable determination. |
conf_force |
Character vector of forced cut expressions. |
conf.cont_medians |
Character vector of continuous confounders to cut at median. |
conf.cont_medians_force |
Character vector of additional continuous confounders to force median cut. |
conf.cont_jcuts |
Named list of positive integers, one per
continuous confounder for which J-quantile cuts are desired. For a
variable
Variables not listed here retain default behaviour. J-quantile
cuts are unconditional (not filtered by LASSO), matching
|
dina_cuts |
Character vector of DINA cut expressions (optional),
merged into the candidate pool exactly like |
collapse_cuts |
Logical. If TRUE, collapse near-redundant continuous
candidate cuts after the full pool is assembled and resolved to literal
numerics. Cuts on the same variable and operator whose thresholds lie
within a per-variable standard-error band are merged to a single
rounded-centroid threshold, subject to a membership safety check; see
|
collapse_cuts_args |
List of overrides for the coarsening, merged onto
the defaults |
replace_med_grf |
Logical. If TRUE, removes median cuts that overlap with GRF cuts. |
defaultcut_names |
Character vector of confounders to force default cuts. |
cut_type |
Character. "default" or "median" for cut strategy. |
exclude_cuts |
Character vector of cut expressions to exclude. |
outcome.name |
Character. Name of outcome variable. |
event.name |
Character. Name of event indicator variable. |
details |
Logical. If TRUE, prints details during execution. |
outcome_type |
Character. One of |
offset.name |
Character or |
Value
A list with components:
dfData frame with derived subgroup factor columns.
confs_namesCharacter vector of factor column names added to
df.confsNamed character vector of cut expressions (continuous cuts plus categorical indicators).
lassokeepCharacter vector of confounders retained by LASSO screening (empty if
use_lasso = FALSE).lassoomitCharacter vector of confounders dropped by LASSO screening (empty if
use_lasso = FALSE).
Examples
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
fs_data <- get_FSdata(df.analysis = df,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", event.name = "status",
use_lasso = FALSE, use_grf = FALSE)
names(fs_data)
Get Best Model from Comparison
Description
Extracts the best fitting model object from a comparison result. If no single best model can be determined, returns the Weibull model if selected by either AIC or BIC. Defaults to Weibull0 model if no model can be determined.
Usage
get_best_survreg(comparison_result)
Arguments
comparison_result |
Output from compare_survreg_models or compare_multiple_survreg |
Value
A survreg model object (defaults to Weibull0 model)
Get all exported functions from ForestSearch namespace
Description
Get all exported functions from ForestSearch namespace
Usage
get_bootstrap_exports()
Get all combinations of subgroup factors up to maxk
Description
Generates all possible combinations of subgroup factors up to a specified maximum size.
Usage
get_combinations_info(L, maxk)
Arguments
L |
Integer. Number of subgroup factors. |
maxk |
Integer. Maximum number of factors in a combination. |
Value
List with max_count (total combinations) and indices_list (indices for each k).
Examples
# L=6 single-factor subgroups, combinations up to 2 factors
info <- get_combinations_info(L = 6, maxk = 2)
info$max_count # total number of combinations
length(info$indices_list)
Get forced cut expressions for variables
Description
For each variable in conf.force.names, returns cut expressions if continuous.
Usage
get_conf_force(df, conf.force.names, cont.cutoff = 4)
Arguments
df |
Data frame. |
conf.force.names |
Character vector of variable names. |
cont.cutoff |
Integer. Cutoff for continuous. |
Value
Character vector of cut expressions.
Examples
df <- data.frame(age = c(45, 60, 35), size = c(20, 35, 15), grade = c(1, 2, 3))
get_conf_force(df, conf.force.names = c("age", "size"))
Get J-quantile cut expressions for variables
Description
For each named variable in conf_jcuts, returns J
deferred cut expressions of the form "X <= qj(X, k, J + 1)"
via cut_var_jq(). Variables that are not continuous (per
cont.cutoff) are skipped with a warning. This is the
cut_var_jq/J-quantile analog of get_conf_force().
Usage
get_conf_force_jq(df, conf_jcuts, cont.cutoff = 4)
Arguments
df |
Data frame. |
conf_jcuts |
Named list of positive integers. Each name is a
continuous-variable column in |
cont.cutoff |
Integer. Cutoff for continuous determination
(passed through to |
Details
Conflicts with other forced-cut mechanisms (defaultcut_names,
conf.cont_medians_force) are validated upstream by
get_FSdata(); this helper performs only structural checks.
Value
Character vector of cut expressions (length sum(J_v)
over continuous v), or character(0) if none.
Examples
df <- data.frame(age = c(45, 60, 35, 50, 70, 25, 80, 55))
# 5 binary cuts at the 1/6, 2/6, ..., 5/6 quantiles of age:
get_conf_force_jq(df, conf_jcuts = list(age = 5))
Get indicator vector for selected subgroup factors
Description
Returns a vector indicating which factors are included in a subgroup combination.
Usage
get_covs_in(
kk,
maxk,
L,
counts_1factor,
index_1factor,
counts_2factor = NULL,
index_2factor = NULL,
counts_3factor = NULL,
index_3factor = NULL
)
Arguments
kk |
Integer. Index of the combination. |
maxk |
Integer. Maximum number of factors in a combination. |
L |
Integer. Number of subgroup factors. |
counts_1factor |
Integer. Number of single-factor combinations. |
index_1factor |
Matrix of indices for single-factor combinations. |
counts_2factor |
Integer. Number of two-factor combinations. |
index_2factor |
Matrix of indices for two-factor combinations. |
counts_3factor |
Integer. Number of three-factor combinations. |
index_3factor |
Matrix of indices for three-factor combinations. |
Value
Numeric vector indicating selected factors (1 = included, 0 = not included).
Examples
# Build index matrices directly (L=4, maxk=2)
L <- 4
idx1 <- matrix(seq_len(L), ncol = 1) # single-factor indices
idx2 <- t(utils::combn(L, 2)) # two-factor indices
# Combination 3 falls in the single-factor block (kk=3 -> factor 3)
get_covs_in(kk = 3, maxk = 2, L = L,
counts_1factor = nrow(idx1),
index_1factor = idx1,
counts_2factor = nrow(idx2),
index_2factor = idx2)
Get variable name from cut expression
Description
Extracts the variable name from a cut expression.
Usage
get_cut_name(thiscut, confounders.name)
Arguments
thiscut |
Character string of the cut expression. |
confounders.name |
Character vector of confounder names. |
Value
Character vector of variable names.
Bootstrap Confidence Interval and Bias Correction Results
Description
Calculates confidence intervals and bias-corrected estimates for bootstrap results.
Usage
get_dfRes(
Hobs,
seHobs,
H1_adj,
H2_adj = NULL,
ystar,
cov_method = "standard",
cov_trim = 0,
est.scale = "hr",
est.loghr = TRUE
)
Arguments
Hobs |
Numeric. Observed estimate. |
seHobs |
Numeric. Standard error of observed estimate. |
H1_adj |
Numeric. Bias-corrected estimate 1. |
H2_adj |
Numeric. Bias-corrected estimate 2 (optional). |
ystar |
Matrix of bootstrap samples. |
cov_method |
Character. Covariance method ("standard" or "nocorrect"). |
cov_trim |
Numeric. Trimming proportion for covariance (default: 0.0). |
est.scale |
Character. "hr" or "1/hr". |
est.loghr |
Logical. Is estimate on log(HR) scale? |
Value
Data.table with confidence intervals and estimates.
Examples
## Not run:
# get_dfRes is called internally by forestsearch_bootstrap_dofuture().
# To use directly, supply bootstrap sample matrix from bootstrap_ystar():
# ystar <- bootstrap_ystar(df, nb_boots = 100, seed = 42L)
# res <- get_dfRes(Hobs = log(1.5), seHobs = 0.3,
# H1_adj = 0.05, H2_adj = 0.03,
# ystar = ystar)
## End(Not run)
Generate Prediction Dataset with Subgroup Treatment Recommendation
Description
Creates a prediction dataset with a treatment recommendation flag based
on the subgroup definition. Supports both label expressions
(e.g., "\{er <= 0\}") and bare column names (e.g., "q3.1").
Usage
get_dfpred(df.predict, sg.harm, version = 1)
Arguments
df.predict |
Data frame for prediction (test or validation set). |
sg.harm |
Character vector of subgroup-defining labels. Values may
be wrapped in braces and optionally negated, e.g. |
version |
Integer; encoding version (maintained for backward compatibility). Default: 1. |
Details
Each element of sg.harm is processed as follows:
Outer braces and leading
!are stripped.If the result matches
"var op value"(whereopis one of<=,<,>=,>,==,!=), the comparison is executed directly ondf.predict[[var]].Otherwise the expression is treated as a column name and membership is
df.predict[[name]] == 1.
Value
Data frame with treatment recommendation flag
(treat.recommend): 0 for harm subgroup, 1 for complement.
See Also
evaluate_comparison for the operator-dispatch
logic, forestsearch for the main analysis function.
Examples
## Not run:
# With brace-wrapped label expressions
sg <- c("{er <= 0}", "{size <= 35}")
df_out <- get_dfpred(df.predict = test_data, sg.harm = sg)
# With negation
sg_neg <- c("{er <= 0}", "!{size <= 35}")
df_neg <- get_dfpred(df.predict = test_data, sg.harm = sg_neg)
# With bare column names (binary indicators)
sg_col <- c("q1.1", "q3.1")
df_col <- get_dfpred(df.predict = encoded_data, sg.harm = sg_col)
## End(Not run)
Extract HR from DGM (Backward Compatible)
Description
Extracts hazard ratios from DGM object, supporting both old and new formats. Also supports CDE (controlled direct effect) extraction for Table 5 of Leon et al. (2024) alignment (theta-ddagger).
Usage
get_dgm_hr(dgm, which = "hr_H")
Arguments
dgm |
DGM object (gbsg_dgm or aft_dgm_flex) |
which |
Character. Which HR to extract: |
Value
Numeric hazard ratio value
Create DGM with Output File Path
Description
Wrapper function that creates a GBSG DGM and generates a standardized output file path for saving results.
Usage
get_dgm_with_output(
model_harm,
n,
k_treat = 1,
target_hr_harm = NULL,
cens_type = "weibull",
out_dir = NULL,
file_prefix = "sim",
file_suffix = "",
include_hr_in_name = FALSE,
verbose = FALSE,
...
)
Arguments
model_harm |
Character. Model type ("alt" or "null") |
n |
Integer. Planned sample size (for filename) |
k_treat |
Numeric. Treatment effect multiplier |
target_hr_harm |
Numeric. Target HR for harm subgroup (used for calibration when model = "alt") |
cens_type |
Character. Censoring type |
out_dir |
Character. Output directory path. If NULL, no file path is generated |
file_prefix |
Character. Prefix for output filename |
file_suffix |
Character. Suffix for output filename |
include_hr_in_name |
Logical. Include achieved HR in filename. Default: FALSE |
verbose |
Logical. Print diagnostic information. Default: FALSE |
... |
Additional arguments passed to |
Value
List with components:
- dgm
The gbsg_dgm object
- out_file
Character path to output file (NULL if out_dir is NULL)
- k_inter
The k_inter value used (either calibrated or default)
Get L_eff at a Given Sample Size
Description
Convenience function: loads calibration and computes
L_{\text{eff}}(N) = C \cdot (N / n_{\min})^\alpha.
Usage
get_leff(N, outcome_type, output_dir = NULL)
Arguments
N |
Integer. Sample size. |
outcome_type |
Character. Outcome type to load. |
output_dir |
Character. Default: NULL (auto). |
Value
Numeric L_eff value.
Get Parameter with Default Fallback
Description
Safely retrieves a named element from a list, returning a default value
if the element is missing or NULL.
Usage
get_param(args_list, param_name, default_value)
Arguments
args_list |
List to extract from. |
param_name |
Character. Name of the element to retrieve. |
default_value |
Default value to return if element is missing or
|
Value
The value of args_list[[param_name]] if present and
non-NULL, otherwise default_value.
Get Subgroup Display Labels for Harm or Benefit Notation
Description
Returns Unicode label components for H/Hc (harm search) or G/Gc (benefit
search) notation. Used internally by build_classification_table,
build_estimation_table, interpret_estimation_table, and
format_oc_results to switch between the two notation systems
described in Leon et al. (2024, Sections 2–4.2).
Usage
get_sg_labels(notation = c("harm", "benefit"))
Arguments
notation |
Character. |
Value
A named list with components:
- sg, sg_c
True subgroup and complement (e.g., "H", "Hc")
- sg_hat, sg_hat_c
Estimated subgroup and complement (Unicode hat)
- plain, plain_c
Plain-text labels for metric names
- word
Descriptive word: "harm" or "benefit"
Fast Cox Model HR Estimation
Description
Fits a minimal Cox model to estimate hazard ratio with reduced overhead. Disables robust variance, model matrix storage, and other extras for speed.
Usage
get_split_hr_fast(df, cox_init = 0, adjust_covariates = NULL)
Arguments
df |
data.frame or data.table with Y, Event, Treat columns. |
cox_init |
Numeric. Initial value for coefficient (default 0).
Ignored when |
adjust_covariates |
Character vector or |
Value
Numeric. Estimated hazard ratio for Treat, or NA if model fails.
Examples
set.seed(42)
df <- data.frame(
Y = rexp(80),
Event = rbinom(80, 1, 0.6),
Treat = rep(0:1, 40)
)
get_split_hr_fast(df)
Get subgroup membership vector
Description
Returns a vector indicating subgroup membership (1 if all selected factors are present, 0 otherwise).
Usage
get_subgroup_membership(zz, covs.in)
Arguments
zz |
Matrix or data frame of subgroup factor indicators. |
covs.in |
Numeric vector indicating which factors are selected (1 = included). |
Value
Numeric vector of subgroup membership (1/0).
Examples
zz <- matrix(c(1, 0, 1, 1, 0, 1), nrow = 3,
dimnames = list(NULL, c("er_le0", "grade3")))
get_subgroup_membership(zz, covs.in = c(1, 1)) # both factors required
get_subgroup_membership(zz, covs.in = c(1, 0)) # first factor only
Target Estimate and Standard Error for Bootstrap
Description
Calculates target estimate and standard error for bootstrap samples.
Usage
get_targetEst(x, ystar, cov_method = c("standard", "nocorrect"), cov_trim = 0)
Arguments
x |
Numeric vector of estimates. |
ystar |
Matrix of bootstrap samples. |
cov_method |
Character. Covariance method ("standard" or "nocorrect"). |
cov_trim |
Numeric. Trimming proportion for covariance (default: 0.0). |
Value
List with target estimate, standard errors, and correction term.
ggplot2 / patchwork forest plot
Description
Creates a publication-quality forest plot using ggplot2 for the CI panel
and patchwork to assemble label and annotation columns alongside it.
Unlike forestploter, fig.height maps directly to row density —
row_height = fig.height / n_rows with no hidden scaling.
Usage
gg_forest(
subgroups,
est,
lo,
hi,
cat_vec = NULL,
cat_colours = NULL,
annot = NULL,
ref_line = 1,
vert_lines = NULL,
ref_col = "firebrick",
ref_lty = "dashed",
vert_col = "grey50",
vert_lty = "dotted",
xlim = NULL,
ticks_at = NULL,
tick_labels = NULL,
xlog = TRUE,
xlab = "Hazard Ratio",
title = NULL,
subtitle = NULL,
footnote = NULL,
point_size = 2.5,
line_size = 0.8,
point_shape = 21,
base_size = 11,
widths = NULL,
row_expand = 0.6,
clip_marker = c("none", "arrow")
)
Arguments
subgroups |
Character vector of subgroup names (displayed top to bottom). |
est |
Numeric vector of point estimates (median HR or similar). |
lo |
Numeric vector of lower bounds (e.g. 1st percentile ECI). |
hi |
Numeric vector of upper bounds (e.g. 99th percentile ECI). |
cat_vec |
Optional character vector of category labels (one per row). Used to colour CI lines and label text. |
cat_colours |
Optional named character vector mapping category labels to colours. Defaults to grey for all rows. |
annot |
Optional named list of character vectors, one per annotation
column. Names become column headers. Each vector must match |
ref_line |
Numeric. X position of the primary reference line (default 1). Drawn as a dashed red line. |
vert_lines |
Numeric vector. X positions of secondary vertical lines (default NULL). Drawn as dotted grey lines. |
ref_col |
Colour of the primary reference line (default "firebrick"). |
ref_lty |
Line type of the primary reference line (default "dashed"). |
vert_col |
Colour of secondary vertical lines (default "grey50"). |
vert_lty |
Line type of secondary vertical lines (default "dotted"). |
xlim |
Numeric vector length 2. X-axis limits for the CI panel. |
ticks_at |
Numeric vector. X-axis tick positions. |
tick_labels |
Character vector. Custom tick labels (default: as.character(ticks_at)). |
xlog |
Logical. If TRUE (default), x-axis on log scale. |
xlab |
Character. X-axis label (default "Hazard Ratio"). |
title |
Character. Overall plot title (default NULL). |
subtitle |
Character. Plot subtitle (default NULL). |
footnote |
Character. Footnote appended below the CI panel (default NULL). |
point_size |
Numeric. Size of point estimate symbol (default 2.5). |
line_size |
Numeric. Line width of CI segments (default 0.8). |
point_shape |
Integer. pch for point estimates (default 21, filled circle). |
base_size |
Numeric. ggplot2 base font size in pt (default 11). Controls all text — increase to make the plot larger; no other knob needed. |
widths |
Numeric vector. Relative patchwork column widths: c(label, ci, annot_1, annot_2, …). Default: c(3.5, 5, rep(1, n_annot)). |
row_expand |
Numeric. Extra space above and below row range on y-axis, in row units (default 0.6). |
clip_marker |
Character. How to mark whiskers whose 1st or 99th
percentile falls outside |
Value
A patchwork object. Render with print() or plot().
Control dimensions entirely via knitr chunk options fig.width /
fig.height: row height = fig.height / n_rows.
Examples
## Not run:
# Recommended fig.height: n_rows * 0.45 + 1.5 (for title/axis overhead)
# e.g. 20 rows -> fig.height = 20 * 0.45 + 1.5 = 10.5
## End(Not run)
Fit GLM with Spline Treatment-Biomarker Interaction
Description
Estimates treatment effects as a function of a continuous covariate using
a generalized linear model with natural cubic splines. This is the GLM
analog of cox_cs_fit for binary, continuous, and count
outcomes.
Usage
glm_cs_fit(
df,
outcome_name = "event",
treat_name = "treat",
z_name = "bm",
family = "binomial",
offset_name = NULL,
alpha = 0.05,
spline_df = 3,
z_by = 0.05,
z_quantile = 0.95,
z_max = Inf,
z_window = 0,
show_plot = FALSE,
plot_params = NULL,
truebeta_name = NULL,
verbose = TRUE
)
## S3 method for class 'glm_cs_fit'
print(x, ...)
Arguments
df |
Data frame containing outcome data. |
outcome_name |
Character. Name of the outcome variable.
Default: |
treat_name |
Character. Name of the treatment variable (0/1).
Default: |
z_name |
Character. Name of the continuous biomarker covariate.
Default: |
family |
Character or family object. One of |
offset_name |
Character or |
alpha |
Numeric. Significance level for confidence intervals (two-sided). Default: 0.05 (95 percent CI). |
spline_df |
Integer. Degrees of freedom for the natural spline. Default: 3. |
z_by |
Numeric. Increment for the biomarker prediction grid. Default: 0.05. |
z_quantile |
Numeric (0–1). Upper quantile for the prediction grid. Default: 0.95 (5th to 95th percentile). |
z_max |
Numeric. Maximum z value for predictions. Default:
|
z_window |
Numeric. Half-width for counting observations near each grid point. Default: 0.0. |
show_plot |
Logical. Display a base-R diagnostic plot.
Default: |
plot_params |
List. Optional plot parameter overrides (see
|
truebeta_name |
Character or |
verbose |
Logical. Print diagnostic information.
Default: |
x |
A |
... |
Additional arguments (unused). |
Details
Model Structure
For binary outcomes (family = "binomial"):
\text{logit}(P(Y=1|Z,A)) = \beta_0 A + f(Z) + g(Z) \cdot A
For count outcomes with offset (family = "poisson"):
\log(E[Y|Z,A]) = \beta_0 A + f(Z) + g(Z) \cdot A + \log(t)
For continuous outcomes (family = "gaussian"):
E[Y|Z,A] = \beta_0 A + f(Z) + g(Z) \cdot A
The treatment-effect profile \beta(Z) = \beta_0 + g(Z) is
estimated on the link scale (log-OR, log-IRR, or identity) with
pointwise delta-method confidence intervals.
Value
A list of class "glm_cs_fit" containing:
- z_profile
Numeric vector. Biomarker grid values.
- loghr_est
Numeric vector. Point estimates of the log treatment effect (log-OR, log-IRR, or MD) at each grid point.
- loghr_lower
Numeric vector. Lower confidence bounds.
- loghr_upper
Numeric vector. Upper confidence bounds.
- se_loghr
Numeric vector. Standard errors (delta method).
- counts_profile
Integer vector. Observation counts near each grid point.
- glm_primary
Numeric. Overall treatment effect from the no-interaction GLM.
- model_fit
The fitted
glmobject.- spline_basis
The natural spline basis object.
- family
Character. The GLM family used.
- effect_label
Character. Human-readable label for the effect scale (e.g., "log(OR)", "log(IRR)", "Mean Difference").
- lrt_pvalue
Numeric. P-value from the likelihood ratio test comparing the interaction model to the main-effects-only model.
- alpha
Numeric. Significance level used.
- ci_level
Numeric. Confidence level (1 - alpha).
Invisibly returns x.
See Also
cox_cs_fit for the survival analog,
glm_effect_profile for the extended interface with
sandwich SEs and overdispersion correction.
Examples
## Not run:
# Binary outcome: log-OR profile over a biomarker
set.seed(42)
df <- data.frame(
event = rbinom(500, 1, 0.3),
treat = rbinom(500, 1, 0.5),
bm = rnorm(500, 0, 1)
)
result <- glm_cs_fit(df, outcome_name = "event", z_name = "bm",
family = "binomial")
# Poisson outcome with offset: log-IRR profile
df$ftime <- rexp(500, 0.01) + 1
result_irr <- glm_cs_fit(df, outcome_name = "event", z_name = "bm",
family = "poisson", offset_name = "ftime")
## End(Not run)
GLM Treatment-Effect Profile for Continuous Biomarker Interaction
Description
Estimates a treatment effect profile across a continuous biomarker using a generalized linear model with natural cubic spline interactions. Supports binary, continuous, and count or rate outcomes. Standard errors are computed via the delta method applied to the full model covariance, with optional sandwich variance estimation for modified-Poisson models.
Usage
glm_effect_profile(
df,
outcome.name,
treat.name = "treat",
z.name = "bm",
effect_measure = c("log_OR", "log_RR", "log_IRR", "RD", "MD"),
offset.name = NULL,
overdispersion = c("none", "quasi", "negbin"),
spline_df = 3L,
z_by = 0.05,
z_quantile = 0.95,
alpha = 0.05,
conf.level = NULL,
strata.name = NULL,
show_plot = FALSE,
verbose = FALSE
)
Arguments
df |
Data frame containing the analysis data. |
outcome.name |
Character. Name of the outcome variable. Binary outcomes must be coded 0/1; count outcomes must be non-negative integers; continuous outcomes may be any numeric. |
treat.name |
Character. Name of the binary treatment indicator
(0 = control, 1 = treated). Default: |
z.name |
Character. Name of the continuous biomarker (effect modifier).
Default: |
effect_measure |
Character. Effect measure to estimate on the link
scale. Allowed values: |
offset.name |
Character or |
overdispersion |
Character. Overdispersion correction for Poisson
family models. |
spline_df |
Integer. Degrees of freedom for natural spline basis.
Default: |
z_by |
Numeric. Increment for the biomarker prediction grid.
Default: |
z_quantile |
Numeric. Upper quantile of the biomarker used as the
grid endpoint (avoids extrapolation into sparse tails).
Default: |
alpha |
Numeric. Two-sided significance level for confidence
intervals. Default: |
conf.level |
Numeric or |
strata.name |
Character or |
show_plot |
Logical. Display a base-R diagnostic plot of the
estimated treatment effect profile. Default: |
verbose |
Logical. Print diagnostic messages. Default: |
Value
An object of class "glm_effect_profile": a named list
containing z_profile (biomarker grid), est (point
estimates on the link scale), lower and upper
(confidence bounds), se (standard errors),
effect_measure, family_used, overdispersion,
dispersion (quasi or negbin dispersion estimate, else NA),
model_fit, spline_basis, and alpha.
See Also
glm_cs_fit for a simpler family-based interface,
cox_cs_fit for survival outcomes,
grf.subg.harm.glm for GRF-based subgroup identification.
Evaluates the performance of GRF-identified subgroups, including hazard ratios, bias, and predictive values. This function is typically used in simulation studies to assess the performance of the GRF subgroup identification method.
Description
Evaluates the performance of GRF-identified subgroups, including hazard ratios, bias, and predictive values. This function is typically used in simulation studies to assess the performance of the GRF subgroup identification method.
Usage
grf.subg.eval(
df,
grf.est,
dgm,
cox.formula.sim,
cox.formula.adj.sim,
analysis = "GRF",
frac.tau = 1
)
Arguments
df |
Data frame containing the analysis data. |
grf.est |
List. Output from |
dgm |
List. Data-generating mechanism (truth) for simulation. |
cox.formula.sim |
Formula for unadjusted Cox model. |
cox.formula.adj.sim |
Formula for adjusted Cox model. |
analysis |
Character. Analysis label (default: "GRF"). |
frac.tau |
Numeric. Fraction of tau for GRF horizon (default: 1.0). |
Value
A data frame with evaluation metrics.
Examples
## Not run:
# grf.subg.eval() is called internally to evaluate GRF subgroup quality.
# See grf.subg.harm.survival() for the standard entry point.
## End(Not run)
GRF Subgroup Identification for GLM Outcomes
Description
Identifies treatment effect subgroups for binary, continuous, and count outcomes using Generalized Random Forests (GRF) for candidate factor screening, followed by exhaustive subgroup enumeration with GLM-based effect estimation.
Usage
grf.subg.harm.glm(
data,
confounders.name,
outcome.name,
treat.name = "treat",
id.name = "id",
outcome_type = c("binary", "continuous", "count"),
effect_measure = NULL,
offset.name = NULL,
overdispersion = c("none", "quasi", "negbin"),
grf_count_transform = c("log", "identity"),
n.min = 60L,
dmin.grf = 0,
frac.tau = 0.5,
maxdepth = 2L,
RCT = TRUE,
sg.criterion = c("mDiff", "Nsg"),
conf.level = 0.95,
seedit = 8316951L,
return_selected_cuts_only = FALSE,
adverse_outcome = FALSE,
tune_grf = FALSE,
grf_selection = c("tree", "frontier"),
frontier_rule = c("effMaxSG", "eff", "maxSG", "minSG", "effMinSG"),
effect_neighborhood = 0.1,
selection_rule = c("neighborhood", "pareto", "both"),
details = FALSE,
verbose = FALSE
)
Arguments
data |
Data frame containing the analysis data. |
confounders.name |
Character vector of potential effect modifier (confounder) names. These are the candidate split variables for GRF and the subgroup enumeration. |
outcome.name |
Character. Name of the outcome column. |
treat.name |
Character. Name of the binary treatment indicator (0/1).
Default: |
id.name |
Character. Name of the subject ID column.
Default: |
outcome_type |
Character. Type of outcome. One of |
effect_measure |
Character. Effect measure to estimate. One of
|
offset.name |
Character or |
overdispersion |
Character. Overdispersion correction for count models.
One of |
grf_count_transform |
Character. Transformation applied to the count
outcome before passing it to |
n.min |
Integer. Minimum subgroup sample size for a valid split.
Default: |
dmin.grf |
Numeric. Minimum absolute CATE magnitude (on the
doubly robust score scale) required for the identified subgroup to
be declared valid. The score equals |
frac.tau |
Numeric. Fraction of the sample used for the GRF horizon
(time-horizon analogue for non-survival outcomes). Default: |
maxdepth |
Integer. Maximum depth for GRF policy trees. Default: |
RCT |
Logical. Whether data come from a randomised trial. When
|
sg.criterion |
Character. Subgroup selection criterion. One of
|
conf.level |
Numeric. Confidence level for subgroup effect estimates.
Default: |
seedit |
Integer. Random seed for GRF. Default: |
return_selected_cuts_only |
Logical. If |
adverse_outcome |
Logical. If |
tune_grf |
Logical. If |
grf_selection |
Character, one of |
frontier_rule |
Character, one of |
effect_neighborhood |
Numeric in (0, 1); relative neighborhood for the
|
selection_rule |
Character, one of |
details |
Logical. Print GRF diagnostic information. Default: |
verbose |
Logical. Print progress messages. Default: |
Details
For count / rate outcomes, the GRF screening step applies a
variance-stabilising \log(Y + 0.5) transformation so that
grf::causal_forest() can be applied without dedicated
count-specific infrastructure. All subgroup effect estimates are
computed from correctly specified Poisson, quasi-Poisson, or
negative-binomial GLMs via create_glm_row.
GRF Screening for Count Outcomes
grf::causal_forest() is designed for continuous outcomes. For count
data, this function transforms Y before passing it to GRF according
to the grf_count_transform argument:
"log"(default)Applies
Y_star = log(Y + 0.5). The +0.5 shift avoidslog(0)for zero-count observations; the log scale approximates variance stabilisation for Poisson data and is recommended when any counts are zero or small."identity"Passes raw counts
Y_star = Ydirectly. Appropriate when counts are large, rarely zero, and the raw scale is near-continuous (e.g., daily hospitalisation volumes in a large centre).
In both cases the transformation applies only to GRF factor
screening and policy-tree cut generation – it determines which
variables and which cut-points are candidate subgroup definitions.
The actual effect estimates in each subgroup are always computed from a
correctly specified Poisson (or quasi-Poisson / negative-binomial) GLM via
create_glm_row, regardless of grf_count_transform.
Alternatively, setting use_grf = FALSE in the parent
forestsearch call routes to LASSO-based factor selection,
which supports family = "poisson" natively via glmnet.
Offset and Person-Time
When offset.name is supplied, \log(\text{exposure}_i) enters
every GLM fit in the subgroup enumeration loop, so that subgroup effects
are on the incidence-rate-ratio scale adjusted for differential follow-up.
Value
A list of class "grf_glm_result" containing:
- sg.harm.id
Character vector of cut expressions defining the identified subgroup (e.g.,
c("v1=1", "v3=1")), length equal to the depth of the selected policy tree.NULLif no subgroup was found. Not a per-subject membership indicator; see the Field naming collision section below.- data
The input data frame with a
treat.recommendcolumn added (0= in harm/questionable subgroup,1= complement).- tree.cuts
Named list of GRF policy-tree split points.
- grf_varimp
Named numeric vector of GRF variable importances.
- outcome.name
Character: name of the outcome column.
- treat.name
Character: name of the treatment column.
- id.name
Character: name of the subject-identifier column.
- confounders.name
Character vector of covariate column names.
- effect_measure
Character: the effect measure used.
- outcome_type
Character: the outcome type.
- overdispersion
Character: the overdispersion correction applied.
- offset.name
Character or
NULL: the offset variable name.- grf_count_transform
Character: the transform argument as matched (
"log"or"identity").- grf_y_transform
Character: transformation actually applied to Y for GRF screening (
"log(Y + 0.5)","identity", or"none").
Field naming collision with forestsearch / subgroup.consistency
The field name sg.harm.id has different semantics on
this object versus on the grp.consistency list returned by
subgroup.consistency (and nested inside the
forestsearch result):
| Object | sg.harm.id contains | Length / type |
grf.subg.harm.glm() result (this function) | character vector of cut expressions | character, length = depth of selected tree |
subgroup.consistency() result | per-subject 0/1 membership indicator | integer, length nrow(df)
|
Practical consequence. On this object, pasting
paste(obj$sg.harm.id, collapse = " & ") correctly renders
the identified subgroup. On a grp.consistency list the same
expression silently concatenates a long 0/1 indicator and produces
output like "0 & 0 & 1 & 0 & ...". When writing code that
must handle both object types, dispatch on class first, or use the
top-level sg.harm field that the forestsearch
result exposes.
This naming collision is a documented CRAN-stable API for v0.1.x and v0.2.x; it is expected to be resolved via a deprecation cycle in a future minor release.
See Also
grf.subg.harm.survival for time-to-event outcomes,
create_glm_row for per-subgroup GLM estimation,
glm_effect_profile for continuous biomarker profiles.
Examples
## Not run:
# Simulate count data with heterogeneous treatment effect
set.seed(123)
n <- 600
z1 <- rbinom(n, 1, 0.4) # binary subgroup factor
df_count <- data.frame(
id = seq_len(n),
treat = rbinom(n, 1, 0.5),
v1 = as.factor(z1),
v2 = as.factor(rbinom(n, 1, 0.5)),
person_time = runif(n, 0.5, 2.0),
age = rnorm(n, 60, 10)
)
# Generate outcome: higher event rate for treated in z1 = 1 subgroup
log_rate <- -1.5 +
0.8 * df_count$treat * (df_count$v1 == 1) -
0.3 * df_count$treat * (df_count$v1 == 0) +
log(df_count$person_time)
df_count$events <- rpois(n, lambda = exp(log_rate))
# Run GRF subgroup identification for count outcome (log transform, default)
grf_glm <- grf.subg.harm.glm(
data = df_count,
confounders.name = c("v1", "v2", "age"),
outcome.name = "events",
treat.name = "treat",
id.name = "id",
outcome_type = "count",
effect_measure = "log_IRR",
offset.name = "person_time",
overdispersion = "quasi",
grf_count_transform = "log", # default: log(Y + 0.5)
n.min = 60,
verbose = TRUE
)
print(grf_glm$sg.harm.id)
# Alternative: identity transform (large counts, no zeros)
grf_glm_id <- grf.subg.harm.glm(
data = df_count,
confounders.name = c("v1", "v2", "age"),
outcome.name = "events",
treat.name = "treat",
id.name = "id",
outcome_type = "count",
effect_measure = "log_IRR",
offset.name = "person_time",
overdispersion = "none",
grf_count_transform = "identity", # raw counts passed to GRF
n.min = 60
)
## End(Not run)
GRF Subgroup Identification for Survival Data
Description
Identifies subgroups with differential treatment effect using generalized random forests (GRF) and policy trees. This function uses causal survival forests to identify heterogeneous treatment effects and policy trees to create interpretable subgroup definitions.
Usage
grf.subg.harm.survival(
data,
confounders.name,
outcome.name,
event.name,
id.name,
treat.name,
frac.tau = 1,
n.min = 60,
dmin.grf = 0,
RCT = TRUE,
details = FALSE,
sg.criterion = "mDiff",
maxdepth = 2,
seedit = 8316951,
return_selected_cuts_only = FALSE,
tune_grf = FALSE,
grf_selection = c("tree", "frontier"),
frontier_rule = c("effMaxSG", "eff", "maxSG", "minSG", "effMinSG"),
effect_neighborhood = 0.1,
selection_rule = c("neighborhood", "pareto", "both")
)
Arguments
data |
Data frame containing the analysis data. |
confounders.name |
Character vector of confounder variable names. |
outcome.name |
Character. Name of outcome variable (e.g., time-to-event). |
event.name |
Character. Name of event indicator variable (0/1). |
id.name |
Character. Name of ID variable. |
treat.name |
Character. Name of treatment group variable (0/1). |
frac.tau |
Numeric. Fraction of tau for GRF horizon (default: 1.0). |
n.min |
Integer. Minimum subgroup size (default: 60). |
dmin.grf |
Numeric. Minimum difference in subgroup mean (default: 0.0). |
RCT |
Logical. Is the data from a randomized controlled trial? (default: TRUE) |
details |
Logical. Print details during execution (default: FALSE). |
sg.criterion |
Character. Subgroup selection criterion ("mDiff" or "Nsg"). |
maxdepth |
Integer. Maximum tree depth (1, 2, or 3; default: 2). |
seedit |
Integer. Random seed (default: 8316951). |
return_selected_cuts_only |
Logical. If TRUE, returns only cuts from the tree
depth that identified the selected subgroup meeting |
tune_grf |
Logical. If |
grf_selection |
Character, one of |
frontier_rule |
Character, one of |
effect_neighborhood |
Numeric in (0, 1); relative neighborhood for the
|
selection_rule |
Character, one of |
Details
The return_selected_cuts_only parameter controls which cuts are returned:
- FALSE (default)
Returns all cuts from all fitted trees (depths 1 to
maxdepth). This provides the full set of candidate splits for downstream exploration and is the original behavior for backward compatibility.- TRUE
Returns only cuts from the tree at the depth that identified the "winning" subgroup meeting the
dmin.grfcriterion. This is useful when you want a focused set of cuts associated with the selected subgroup, reducing noise from non-selected trees.
When return_selected_cuts_only = TRUE and no subgroup meets the criteria,
tree.cuts will be empty (character(0)).
Value
A list with GRF results, including:
data |
Original data with added treatment recommendation flags |
grf.gsub |
Selected subgroup information |
sg.harm.id |
Character vector of cut expressions
defining the identified subgroup (length = depth of the selected
policy tree), or |
tree.cuts |
Cut expressions - either all cuts from all trees (if
|
tree.names |
Unique variable names used in cuts |
tree |
Selected policy tree object |
tau.rmst |
Time horizon used for RMST |
harm.any |
All subgroups with positive treatment effect difference |
selected_depth |
Depth of the tree that identified the subgroup (when found) |
return_selected_cuts_only |
Logical indicating which cut extraction mode was used |
Additional tree-specific cuts and objects (tree1, tree2, tree3) based on maxdepth
Field naming collision with forestsearch / subgroup.consistency
The field name sg.harm.id has different semantics on
this object versus on the grp.consistency list returned by
subgroup.consistency (and nested inside the
forestsearch result):
| Object | sg.harm.id contains | Length / type |
grf.subg.harm.survival() result (this function) | character vector of cut expressions | character, length = depth of selected tree |
subgroup.consistency() result | per-subject 0/1 membership indicator | integer, length nrow(data)
|
Practical consequence. On this object, pasting
paste(obj$sg.harm.id, collapse = " & ") correctly renders
the identified subgroup. On a grp.consistency list the same
expression silently concatenates a long 0/1 indicator and produces
output like "0 & 0 & 1 & 0 & ...". When writing code that
must handle both object types, dispatch on class first, or use the
top-level sg.harm field that the forestsearch
result exposes.
This naming collision is a documented CRAN-stable API for v0.1.x and v0.2.x; it is expected to be resolved via a deprecation cycle in a future minor release.
Examples
## Not run:
# Return all cuts (default behavior)
result_all <- grf.subg.harm.survival(
data = trial_data,
confounders.name = c("age", "biomarker", "region"),
outcome.name = "tte",
event.name = "event",
id.name = "id",
treat.name = "treat",
dmin.grf = 0.1,
maxdepth = 2
)
result_all$tree.cuts
# Returns cuts from both depth 1 and depth 2 trees
# Return only cuts from the selected tree
result_selected <- grf.subg.harm.survival(
data = trial_data,
confounders.name = c("age", "biomarker", "region"),
outcome.name = "tte",
event.name = "event",
id.name = "id",
treat.name = "treat",
dmin.grf = 0.1,
maxdepth = 2,
return_selected_cuts_only = TRUE
)
result_selected$tree.cuts
# Returns cuts only from the depth that identified the winning subgroup
## End(Not run)
Cross-validated choice of the shrinkage tuning parameter r
Description
Implements Algorithm 2 of Guo and He (2021), Section 2.5: a v-fold
cross-validated choice of the shrinkage tuning parameter r used by
guohe_algorithm3(). For each candidate r and each fold, the
bias-reduced estimator is computed on the training folds and compared, on
the held-out fold, against every candidate subgroup's effect estimate with
its sampling variance subtracted; the r minimising the resulting CV
objective is selected.
Usage
guohe_adaptive_r(
data,
outcome = c("survival", "binary", "continuous"),
treatment,
candidates,
time = NULL,
event = NULL,
y = NULL,
orient = -1,
r_grid = c(0.03, 0.1, 0.2, 0.3, 0.4, 0.45),
v = 5L,
B = 200L,
level = 0.05,
seed = NULL,
min_events = 5L,
refit = TRUE,
adjust_covariates = NULL,
fast = NULL,
parallel = FALSE
)
Arguments
data |
A data frame. |
outcome |
One of |
treatment |
Name of the 0/1 treatment column. |
candidates |
Character vector of names of 0/1 membership columns forming the enumerated family. |
time, event |
Survival only: names of the time and event columns. |
y |
Binary or continuous only: name of the outcome column. |
orient |
Either |
r_grid |
Numeric vector of candidate values in (0, 0.5). |
v |
Number of cross-validation folds. |
B |
Number of bootstrap resamples. |
level |
One-sided level; the two-sided interval has coverage |
seed |
Optional integer for reproducible resampling. |
min_events |
Minimum events (survival) or minimum per-arm count for a candidate to be estimable. |
refit |
Logical; if |
adjust_covariates |
Optional character vector of covariate names for the
within-subgroup model, matching the argument of the same name in
NOTE: |
fast |
Logical or |
parallel |
Logical; if |
Details
This is an independent implementation of the published method, validated by reproducing the simulation study of Section 5 of Guo and He (2021).
Value
An object of class "guohe_ar": the selected r (r_hat), the CV
objective over the grid (objective), the per-candidate objective at the
selected r (per_candidate), bookkeeping fields (r_grid, v, B,
n, n_candidates, fast, outcome, orient, level,
adjust_covariates), and, when refit = TRUE, the refitted full-data
guohe_algorithm3() result (fit).
References
Guo, X. and He, X. (2021). Inference on Selected Subgroups in Clinical Trials. Journal of the American Statistical Association, 116(535), 1498–1506. doi:10.1080/01621459.2020.1740096
Examples
set.seed(1)
n <- 160
df <- data.frame(
treat = rbinom(n, 1, 0.5),
y = rbinom(n, 1, 0.3),
g1 = rbinom(n, 1, 0.5),
g2 = rbinom(n, 1, 0.5),
g3 = rbinom(n, 1, 0.5)
)
cv <- guohe_adaptive_r(
df, outcome = "binary", treatment = "treat",
candidates = c("g1", "g2", "g3"), y = "y",
orient = +1, r_grid = c(0.05, 0.25), v = 2, B = 25, seed = 2
)
cv$r_hat
De-biased inference for the best subgroup over a large enumerated family
Description
Finite realization of Algorithm 3 of Guo and He (2021): the candidate family is supplied as materialized 0/1 membership columns, the most extreme oriented effect is selected, and its selection optimism is removed by a pair (case) bootstrap with deterministic shrinkage offsets computed against the global maximum. Survival, binary, and continuous outcomes are supported.
Usage
guohe_algorithm3(
data,
outcome = c("survival", "binary", "continuous"),
treatment,
candidates,
time = NULL,
event = NULL,
y = NULL,
orient = -1,
B = 2000L,
r = 0.03,
level = 0.05,
seed = NULL,
min_events = 5L,
diagnostics = TRUE,
parallel = FALSE,
adjust_covariates = NULL,
fast = NULL
)
Arguments
data |
A data frame. |
outcome |
One of |
treatment |
Name of the 0/1 treatment column. |
candidates |
Character vector of names of 0/1 membership columns forming the enumerated family. |
time, event |
Survival only: names of the time and event columns. |
y |
Binary or continuous only: name of the outcome column. |
orient |
Either |
B |
Number of bootstrap resamples. |
r |
Shrinkage tuning parameter, strictly between 0 and 0.5. Guo and He's Algorithm 2 gives a cross-validated choice; a fixed small value is used here. |
level |
One-sided level; the two-sided interval has coverage |
seed |
Optional integer for reproducible resampling. |
min_events |
Minimum events (survival) or minimum per-arm count for a candidate to be estimable. |
diagnostics |
Logical; if |
parallel |
Logical; if |
adjust_covariates |
Optional character vector of covariate names for the
within-subgroup model, matching the argument of the same name in
NOTE: |
fast |
Logical or |
Details
Membership is computed once on the full data and carried through the bootstrap; cut points are NOT re-derived per resample. This is required by the method (the index set is held fixed) and is what makes the comparison against multiplier resampling like-for-like; regenerating the family per resample is the full-bootstrap object instead.
The oriented score is orient * b, with b the within-subgroup treatment
coefficient on the natural scale. orient = -1 treats a protective ratio
(HR or OR below 1) as the effect of interest, matching Guo and He's MONET1
convention; orient = +1 treats a harmful ratio as the effect of interest,
which is the setting for a forest-search harm subgroup.
This is an independent implementation of the published method, validated by reproducing the simulation study of Section 5 of Guo and He (2021).
Value
An object of class "guohe_a3": a list with the selected candidate
(selected), the naive and de-biased effects on the reporting scale
(naive, debiased), the one-sided bound of Guo and He's Algorithm 3 (bound_one_sided,
with its conservative side in bound_side), the two-sided
quantile-inversion interval (ci_two_sided), the heuristic Wald interval
(ci_wald), and bookkeeping fields (n, B, B_requested, r,
level, n_candidates, n_estimable, fast, lean_cox,
adjust_covariates, and, when diagnostics = TRUE, the per-candidate
resample drop_rate).
References
Guo, X. and He, X. (2021). Inference on Selected Subgroups in Clinical Trials. Journal of the American Statistical Association, 116(535), 1498–1506. doi:10.1080/01621459.2020.1740096
Examples
set.seed(1)
n <- 160
df <- data.frame(
treat = rbinom(n, 1, 0.5),
y = rbinom(n, 1, 0.3),
g1 = rbinom(n, 1, 0.5),
g2 = rbinom(n, 1, 0.5),
g3 = rbinom(n, 1, 0.5)
)
fit <- guohe_algorithm3(
df, outcome = "binary", treatment = "treat",
candidates = c("g1", "g2", "g3"), y = "y",
orient = +1, B = 50, seed = 2
)
fit
Check if Matrix Has Positive Variance
Description
Check if Matrix Has Positive Variance
Usage
has_positive_variance(x)
Format Hazard Ratio and Confidence Interval
Description
Formats a hazard ratio and confidence interval for display.
Usage
hrCI_format(hrest)
Arguments
hrest |
Numeric vector with HR, lower, and upper confidence limits. |
Value
Character string formatted as \"HR (lower, upper)\".
Examples
hrCI_format(c(1.45, 0.98, 2.14))
Build an aft_dgm_flex with synthetic regions and paper-style effect structure.
Description
Build an aft_dgm_flex with synthetic regions and paper-style effect structure.
Usage
inject_mrct_structure(
seed_data,
outcome_var,
event_var,
treatment_var,
continuous_vars,
factor_vars,
x_pred,
spline_spec,
region = list(),
x3 = NULL,
region_treat_loghr = NULL,
n_super = 5000L,
expand = c("copula", "bootstrap"),
continuous_vars_cens = NULL,
factor_vars_cens = NULL,
cens_type = "weibull",
seed = 8316951L,
verbose = FALSE
)
Arguments
seed_data |
seed trial data.frame (one row per patient) |
outcome_var, event_var, treatment_var |
column names in seed_data |
continuous_vars, factor_vars |
covariates entering the AFT model (factor_vars must be 0/1 numeric or 2-level factors for clean expansion) |
x_pred |
region-associated effect modifier (X2), a continuous_var |
spline_spec |
list(knot, zeta, log_hrs) on the NATURAL scale of x_pred:
true log-HR is piecewise linear through (0, |
region |
list(name = "region", prevalence = 0.2, or_pred = 10,
x1_vars = character(0), or_x1 = 1, loghr = 0, scale = "rank") – loghr is
the prognostic log-HR of Region (no TE modification; 0 for the paper's
DGM; -log(5) reproduces the Feb-2026 deck's strongly prognostic region);
scale = "rank" (ECDF) or "minmax" (the paper's |
x3 |
NULL, or list(vars = c(...), cuts = list(...), loghr = ...) for a balanced effect modifier: cuts as in generate_aft_dgm_flex (use fixed values), loghr = log-HR change (treatment x subgroup). |
region_treat_loghr |
NULL, or a log-HR for a direct treat x Region interaction (the residual Region-U-TE pathway). Mutually exclusive with x3 (the package supports one interaction term). |
n_super |
size of the expanded super-population |
expand |
"copula" or "bootstrap" |
continuous_vars_cens, factor_vars_cens, cens_type |
passed through |
seed |
RNG seed |
verbose |
logical |
Value
aft_dgm_flex with df_super = expanded population (+ z_<region>), and
a $mrct element documenting the injected structure.
Generate Narrative Interpretation of Estimation Properties
Description
Produces a templated text summary of the estimation properties table, automatically populating numerical results from the simulation output. Useful for reproducible vignettes where interpretation paragraphs should update when simulations are re-run.
Usage
interpret_estimation_table(
results,
dgm,
analysis_method = "FSlg",
n_sims = NULL,
n_boots = 300,
digits = 2,
scenario = NULL,
cat = TRUE,
subgroup_notation = c("harm", "benefit"),
trim_threshold = 1000,
trim_fraction = 0.01
)
Arguments
results |
Data frame of simulation results (same as for
|
dgm |
DGM object with true parameter values. |
analysis_method |
Character. Which analysis method to summarise.
Default: |
n_sims |
Integer. Total number of simulations (for detection rate).
If |
n_boots |
Integer. Number of bootstraps (for narrative). Default: 300. |
digits |
Integer. Decimal places for reported values. Default: 2. |
scenario |
Character. One of
If |
cat |
Logical. If |
subgroup_notation |
Character. |
trim_threshold |
Numeric or |
trim_fraction |
Numeric in |
Value
Invisibly returns the interpretation as a character string.
See Also
build_estimation_table,
format_oc_results, get_dgm_hr
Examples
## Not run:
# In a vignette chunk with results = "asis":
interpret_estimation_table(results_null, dgm_null, scenario = "null")
# Capture for further processing:
txt <- interpret_estimation_table(results_alt, dgm_alt, cat = FALSE)
## End(Not run)
Interpret and Diagnose Search Configuration
Description
Generates a human-readable diagnostic of how ForestSearch will
interpret the user's parameter combination. Prints via
message() and returns an invisible list of diagnostics.
Usage
interpret_search_config(
outcome_type,
effect_measure,
adverse_outcome,
effect_threshold,
consistency_threshold,
use_lasso,
use_grf,
outcome.name,
event.name,
treat.name,
offset.name = NULL,
subgroup_method = "consistency",
quiet = FALSE
)
Arguments
outcome_type |
Character. One of |
effect_measure |
Character. |
adverse_outcome |
Logical. As resolved by forestsearch(). |
effect_threshold |
Numeric. Resolved screening threshold (on log scale for ratio measures, identity for others). |
consistency_threshold |
Numeric. Resolved consistency threshold. |
use_lasso |
Logical. |
use_grf |
Logical. Whether GRF candidate-cut generation is on;
see |
outcome.name |
Character. Name of outcome column. |
event.name |
Character. Name of event column. |
treat.name |
Character. Name of treatment column. |
offset.name |
Character or NULL. |
subgroup_method |
Character, the resolved |
quiet |
Logical. If TRUE, suppress output. |
Value
Invisible list with diagnostic fields.
Invert DGM Effect Values for Benefit Search
Description
Transforms a DGM object's truth values from the switched treatment scale
to the original scale. For ratio-scale measures (HR, OR, IRR), fields
are reciprocated (1/x); for identity-scale measures (MD, RD, IRD),
fields are negated (-x). The effect measure is auto-detected from
dgm$effect_measure when not explicitly provided.
Usage
invert_dgm_hrs(dgm, effect_measure = NULL)
Arguments
dgm |
A DGM object ( |
effect_measure |
Character or |
Details
Called automatically by table functions when
subgroup_notation = "benefit".
Value
Modified DGM object with all effect fields inverted. Non-effect fields (super-population data, model parameters, etc.) are unchanged.
Invert Effect Columns in Simulation Results for Benefit Search
Description
Transforms simulation results from the switched treatment scale to the
original scale. For ratio-scale measures (HR, OR, IRR), inversion is
1/x; for identity-scale measures (MD, RD, IRD), inversion is
-x. Called automatically by table functions when
subgroup_notation = "benefit".
Usage
invert_hr_columns(res, effect_measure = NULL)
Arguments
res |
Data frame of simulation results (from
|
effect_measure |
Character or |
Value
Data frame with all effect-related columns inverted. Non-effect columns (classification metrics, subgroup sizes, etc.) are unchanged.
Check if a variable is continuous
Description
Determines if a variable is continuous based on the number of unique values.
Usage
is.continuous(x, cutoff = 4)
Arguments
x |
A vector. |
cutoff |
Integer. Minimum number of unique values to be considered continuous. |
Value
1 if continuous, 2 if not.
Check if cut expression is for a continuous variable (OPTIMIZED)
Description
Determines if a cut expression refers to a continuous variable. This optimized version avoids redundant lookups by using word boundary matching instead of partial string matching.
Usage
is_flag_continuous(thiscut, confounders.name, df, cont.cutoff)
Arguments
thiscut |
Character string of the cut expression. |
confounders.name |
Character vector of confounder names. |
df |
Data frame. |
cont.cutoff |
Integer. Cutoff for continuous. |
Value
Logical; TRUE if continuous, FALSE otherwise.
Check if cut expression should be dropped
Description
Determines if a cut expression should be dropped (e.g., variable has <=1 unique value).
Usage
is_flag_drop(thiscut, confounders.name, df)
Arguments
thiscut |
Character string of the cut expression. |
confounders.name |
Character vector of confounder names. |
df |
Data frame. |
Value
Logical; TRUE if should be dropped, FALSE otherwise.
Is this Effect Measure on the Identity Scale?
Description
Identity-scale measures (MD, RD, IRD) use additive differences;
ratio-scale measures (HR, OR, RR, IRR) use multiplicative ratios.
The distinction matters for benefit-search inversion: ratio-scale
values are reciprocated (1/x), identity-scale values are
negated (-x).
Usage
is_identity_scale(effect_measure)
Arguments
effect_measure |
Character. One of |
Value
Logical. TRUE for identity-scale measures.
KM median summary for subgroup
Description
Calculates median survival for each treatment group using Kaplan-Meier.
Usage
km_summary(Y, E, Treat)
Arguments
Y |
Numeric vector of outcome. |
E |
Numeric vector of event indicators. |
Treat |
Numeric vector of treatment indicators. |
Value
Numeric vector of medians.
LASSO selection for Cox model
Description
Performs LASSO variable selection using Cox regression.
Usage
lasso_selection(
df,
confounders.name,
outcome.name,
event.name,
seedit = 8316951,
outcome_type = "survival",
offset.name = NULL
)
Arguments
df |
Data frame. |
confounders.name |
Character vector of confounder names. |
outcome.name |
Character. Name of outcome variable. |
event.name |
Character. Name of event indicator variable. |
seedit |
Integer. Random seed. |
outcome_type |
Character. One of |
offset.name |
Character or |
Value
List with selected, omitted variables, coefficients, lambda, and fits.
Examples
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
lasso_selection(df,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", event.name = "status")
Find the Most Recent .rds File Matching a Pattern
Description
Scans output_dir for files matching the given prefix
pattern and returns the path to the most recently modified one.
Date-stamped filenames (e.g., results_20260402.rds) sort
naturally, but this function uses filesystem modification time
for robustness.
Usage
latest_rds(prefix, output_dir = "_data", verbose = TRUE)
Arguments
prefix |
Character. Filename prefix to match
(e.g., |
output_dir |
Character. Directory to scan.
Default: |
verbose |
Logical. Print the matched file. Default: TRUE. |
Value
Full path to the most recent matching file, or NULL if no match is found.
Examples
## Not run:
path <- latest_rds("selection_survival_frontier")
res <- readRDS(path)
path <- latest_rds("calibration_binary_leff_grid")
## End(Not run)
Load L_eff Calibration Results
Description
Reads a calibration bundle saved by save_leff_calibration().
Prints a summary and returns the full bundle.
Usage
load_leff_calibration(outcome_type, output_dir = NULL, verbose = TRUE)
Arguments
outcome_type |
Character. One of "binary", "survival", "count", "continuous". |
output_dir |
Character. Directory containing .rds files.
Default: |
verbose |
Logical. Print summary on load. Default: TRUE. |
Value
A list with components C, alpha, n_min, outcome_type, cal_data, dgm_description, timestamp, etc.
Examples
## Not run:
cal <- load_leff_calibration("binary")
C_prior <- cal$C
alpha_prior <- cal$alpha
## End(Not run)
Make an Effect-Estimator Closure
Description
Returns a closure function(data_slice) that estimates the
within-subgroup treatment effect for the requested outcome type and
effect measure. The closure is called repeatedly inside the
splitting-consistency loop and the bootstrap worker, so it must be
lightweight and handle small-sample edge cases gracefully.
Usage
make_effect_estimator(
outcome_type,
treat.name,
outcome.name,
event.name = NULL,
offset.name = NULL,
effect_measure = NULL,
adverse_outcome = TRUE,
robust_se = TRUE,
adjust_covariates = NULL,
ps_adjust_method = c("none", "iptw", "dr_gcomp"),
...
)
Arguments
outcome_type |
Character. One of |
treat.name |
Character. Name of the binary treatment column (0/1). |
outcome.name |
Character. Name of the outcome column. |
event.name |
Character or |
offset.name |
Character or |
effect_measure |
Character or |
adverse_outcome |
Logical. If |
robust_se |
Logical. Use sandwich robust SE when available.
Default |
adjust_covariates |
Character vector or |
ps_adjust_method |
Character. Propensity score adjustment method.
One of |
... |
Additional arguments (reserved for future extensions). |
Value
A closure function(data_slice) returning a list with
components: estimate, se, converged, n0,
n1, measure, method_used.
Examples
## Not run:
# Binary outcome with odds ratio
fn <- make_effect_estimator("binary", "treat", "event",
effect_measure = "OR")
result <- fn(my_data)
exp(result$estimate) # odds ratio
# Poisson rate with IPTW adjustment
fn_iptw <- make_effect_estimator("binary", "treat", "event",
offset.name = "time", effect_measure = "IRR",
ps_adjust_method = "iptw")
## End(Not run)
Build the logistic region model and draw Region on a population.
Description
Build the logistic region model and draw Region on a population.
Usage
make_region_model(
pop,
x_pred = NULL,
or_pred = 1,
x1_vars = character(0),
or_x1 = 1,
prevalence = 0.2,
scale_ref = NULL,
scale = c("rank", "minmax"),
band = NULL
)
Arguments
pop |
data.frame holding x_pred and x1_vars (natural scale) |
x_pred |
name of the region-associated effect modifier (X2); NULL for no X2 term (regional imbalance only through x1_vars) |
or_pred |
odds ratio for s(x_pred) -> Region (OR = 1: none) |
x1_vars |
names of imbalanced-only covariates (X1); character(0) for none |
or_x1 |
odds ratios for the X1 terms (recycled) |
prevalence |
target marginal P(Region = 1) |
scale_ref |
data.frame used to define the scaling (defaults to pop) |
scale |
"rank" (ECDF, default) or "minmax" (the paper's scaling) |
band |
NULL (default, no band term), or list(var, cut, or): adds
|
Value
list(alpha0, coefs, band, prob = function(df), draw = function(df, seed))
Check Event Count Criteria
Description
Check Event Count Criteria
Usage
meets_event_criteria(event_counts, d0.min, d1.min)
Check Prevalence Threshold
Description
Check Prevalence Threshold
Usage
meets_prevalence_threshold(x, minp)
Multiplier-resampling (MR) estimates table (bootstrap-table format)
Description
Renders the selected subgroup's MR de-biased estimate (forestsearch() with
mr_inference = TRUE) in the same gt layout as the bias-corrected table
produced by summarize_bootstrap_results(). The descriptive columns
(sample size, medians/RMST or rates) and the naive effect do not depend on
which de-biasing method was used, so they are inherited from a full-bootstrap
(FB) result used as a template; the bias-corrected column is replaced with
the MR values from fs$mr_inference – the de-biased estimate for the
selected subgroup and, when MR was run with include_complement = TRUE, for
the complement as well (otherwise the complement cell is marked, since MR
de-biases the selected subgroup only).
Usage
mr_estimates_table(fs, boot_results, est.scale = "hr", digits = 2)
Arguments
fs |
A |
boot_results |
A |
est.scale |
Character, |
digits |
Integer display precision (default |
Details
The selected-subgroup row is matched by its bias-corrected string
(boot_results$SG_CIs$H_bc) rather than by position, so the substitution is
robust to row ordering. The MR de-biased confidence interval is a
leading-order approximation of the full-bootstrap infinitesimal-jackknife
interval – it comes from the IJ variance computed on the multiplier draws
under the default ci_method = "ij", and from the subgroup robust SE only
under ci_method = "wald". Either way the bias-corrected cell may differ
slightly from the full bootstrap even when the naive effect matches exactly.
Value
A gt table.
See Also
forestsearch() for the mr_inference switch and the vocabulary
section; fs_mr_inference() for the MR computation itself;
forestsearch_bootstrap_dofuture() for the full bootstrap that supplies
the template; fs_fdr_report() for harm-confirmation threshold sweeps;
summarize_bootstrap_results(), format_bootstrap_table()
MRCT Regional Subgroup Simulation
Description
Simulates multi-regional clinical trials and evaluates ForestSearch subgroup identification. Splits data by region into training and testing populations, identifies subgroups using ForestSearch on training data, and evaluates performance on the testing region.
Usage
mrct_region_sims(
dgm,
n_sims,
n_sample = NULL,
region_var = "z_regA",
sg_focus = "minSG",
maxk = 1,
hr.threshold = 0.9,
hr.consistency = 0.8,
pconsistency.threshold = 0.9,
confounders.name = NULL,
conf_force = NULL,
fs_args = list(),
sim_args = list(rand_ratio = 1, draw_treatment = TRUE),
analysis_time = 60,
cens_adjust = 0,
parallel_args = list(plan = "multisession", workers = NULL, show_message = TRUE),
details = FALSE,
verbose_n_sims = 2L,
seed = NULL
)
Arguments
dgm |
Data generating mechanism object from |
n_sims |
Integer. Number of simulations to run |
n_sample |
Integer. Sample size per simulation. If NULL (default), uses the entire super-population from dgm |
region_var |
Character. Name of the region indicator variable used to split data into training (region_var == 0) and testing (region_var == 1) populations. Default: "z_regA" |
sg_focus |
Character. Subgroup selection criterion passed to
|
maxk |
Integer. Maximum number of factors in subgroup combinations (1 or 2). Default: 1 |
hr.threshold |
Numeric. |
hr.consistency |
Numeric. |
pconsistency.threshold |
Numeric. |
confounders.name |
Character vector. Confounder variable names for ForestSearch. If NULL, automatically extracted from dgm |
conf_force |
Character vector. Forced cuts to consider in ForestSearch. Default: c("z_age <= 65", "z_bm <= 0", "z_bm <= 1", "z_bm <= 2", "z_bm <= 5") |
fs_args |
Named list. Additional arguments passed directly to
|
sim_args |
Named list. Additional arguments passed to
|
analysis_time |
Numeric. Time of analysis for administrative censoring. Default: 60 |
cens_adjust |
Numeric. Adjustment factor for censoring rate on log scale. Default: 0 |
parallel_args |
List. Parallel processing configuration with components:
|
details |
Logical. Print detailed progress information. Default: FALSE |
verbose_n_sims |
Integer. When |
seed |
Integer. Base random seed for reproducibility. Default: NULL |
Details
Simulation Process
For each simulation:
Sample from super-population using
simulate_from_dgmSplit by region_var into training and testing populations
Estimate HRs in ITT, training, and testing populations
Run
forestsearchon training populationApply identified subgroup to testing population
Calculate subgroup-specific estimates
Region Variable
The region_var parameter is used ONLY for splitting data into training/testing
populations. It does not imply any prognostic effect. To include prognostic
confounder effects, specify them when creating the DGM using
create_dgm_for_mrct or generate_aft_dgm_flex.
Value
A data.table with simulation results containing:
- sim
Simulation index
- n_itt
ITT sample size
- hr_itt
ITT hazard ratio (stratified if strat variable present)
- hr_ittX
ITT hazard ratio stratified by region
- n_train
Training (non-region A) sample size
- hr_train
Training population hazard ratio
- n_test
Testing (region A) sample size
- hr_test
Testing population hazard ratio
- any_found
Indicator: 1 if subgroup identified, 0 otherwise
- sg_found
Character description of identified subgroup
- n_sg
Subgroup sample size
- hr_sg
Subgroup hazard ratio in testing population
- POhr_sg
Potential outcome hazard ratio in subgroup (testing)
- prev_sg
Subgroup prevalence (proportion of testing population)
- n_sg_train
Subgroup sample size in training population
- hr_sg_train
Subgroup hazard ratio in training population
- POhr_sg_train
Potential outcome hazard ratio in subgroup (training)
- hr_sg_null
Subgroup HR when found, NA otherwise
See Also
forestsearch for subgroup identification algorithm
generate_aft_dgm_flex for DGM creation
simulate_from_dgm for data simulation
create_dgm_for_mrct for MRCT-specific DGM wrapper
summaryout_mrct for summarizing simulation results
Examples
## Not run:
# Create DGM for alternative hypothesis
dgm_alt <- create_dgm_for_mrct(
df_case = df_case,
model_type = "alt",
log_hrs = log(c(3, 1.25, 0.50)),
verbose = TRUE
)
# Run simulations
results <- mrct_region_sims(
dgm = dgm_alt,
n_sims = 100,
region_var = "z_regA",
sg_focus = "minSG",
parallel_args = list(plan = "multisession", workers = 4),
details = TRUE
)
# Summarize results
cat("Subgroup identification rate:", mean(results$any_found) * 100, "%\n")
## End(Not run)
Truth by region on the super-population: prevalence, x_pred shift, causal AHR / CDE, and a large-sample Cox HR on UNCENSORED potential outcomes. Compare HR_cox_PO with the censored-trial HR from sim_region_metrics(): a gap between them isolates the prognostic-shift / follow-up-truncation mechanism (the Feb-2026 deck) from covariate shift (AHR itself differing by region).
Description
Truth by region on the super-population: prevalence, x_pred shift, causal AHR / CDE, and a large-sample Cox HR on UNCENSORED potential outcomes. Compare HR_cox_PO with the censored-trial HR from sim_region_metrics(): a gap between them isolates the prognostic-shift / follow-up-truncation mechanism (the Feb-2026 deck) from covariate shift (AHR itself differing by region).
Usage
mrct_truth_by_region(dgm, seed = 1L)
Arguments
dgm |
An MRCT data-generating mechanism returned by
|
seed |
Integer seed passed to |
Value
A data frame with one row per stratum ("ALL", "NON-REGION",
"REGION") and columns stratum, n, prev, x_pred_mean, AHR,
CDE and HR_cox_PO. Attribute "smd_x_pred" holds the standardised
mean difference of the effect modifier between region and non-region.
Calculate n and percent
Description
Returns count and percent for a vector relative to a denominator.
Usage
n_pcnt(x, denom)
Arguments
x |
Vector of values. |
denom |
Denominator for percent calculation. |
Value
Character string formatted as \"n (percent%)\".
Examples
n_pcnt(1:30, 100)
Pareto Dominance Flags
Description
Internal helper. For each row of result_dt, returns
TRUE if the row is dominated in (hr, N) space
(both maximized) by another row, and FALSE otherwise. Rows
with NA hr or N are flagged as dominated.
Usage
pareto_dominated_flags(result_dt, effect_log_scale = FALSE)
Arguments
result_dt |
Data.table of candidate subgroups with columns
|
effect_log_scale |
Logical. If |
Details
Used by both compute_pareto_frontier() (which filters by
the negation of this vector) and the selection_rule = "pareto"
branches of sort_subgroups() / sort_subgroups_preview()
(which use it as an inclusion indicator).
Value
Logical vector of length nrow(result_dt).
Format Pareto Frontier of Candidate Subgroups
Description
Renders the post-hoc Pareto frontier on (effect, N) – both maximized
– as a formatted gt table or returns it as a
data.table for programmatic use. Works uniformly across
survival (Cox PH) and GLM outcome types: the effect-column label and
scale handling are derived from the forestsearch object's
effect_measure.
Usage
pareto_frontier_table(
fs,
format = c("gt", "data.table"),
digits = 2L,
digits_effect = NULL,
digits_pcons = NULL,
digits_ci = NULL,
include_dominated = FALSE,
include_factor_columns = TRUE,
highlight_selected = TRUE,
ci_table = NULL
)
Arguments
fs |
A |
format |
Character. Either |
digits |
Integer. Master decimal-places setting that drives
all three column-specific defaults ( |
digits_effect |
Integer or |
digits_pcons |
Integer or |
digits_ci |
Integer or |
include_dominated |
Logical. If |
include_factor_columns |
Logical. If |
highlight_selected |
Logical. If |
ci_table |
Optional data.table of frontier CIs from
|
Details
The frontier is a diagnostic: it lists candidate subgroups that are
not dominated on (effect, N) simultaneously. It is computed
inside sg_consistency_out (see
compute_pareto_frontier) and attached to the result
object as fs$grp.consistency$out_sg$pareto_frontier. It is
not used for selection – the selected subgroup is chosen
by sg_focus and may or may not appear on the frontier.
Scale handling. For ratio measures stored on the log
scale internally ("OR", "RR", "IRR"), the
effect column is exponentiated for display. For "HR" (Cox
PH, natural scale) and identity-scale measures ("RD",
"IRD", "MD"), values pass through unchanged.
Selected-row identification. The selected subgroup is
identified by matching the original-table row index m
against the top row of fs$grp.consistency$out_sg$result.
This is robust to sorting and outcome-type differences.
For sg_focus = "hrMinSG", the selected subgroup may
not appear on the frontier – that focus deliberately prefers small
subgroups, which are typically N-dominated.
Optional confidence intervals. When ci_table is
supplied (the data.table returned by
compute_frontier_cis), three CI columns are added to
the table. All three CIs are computed by
compute_frontier_cis; this function only displays
what is passed in. Pass the same ci_table to
plot_pareto_frontier for consistent display across
table and plot.
-
Naive 95% CI– full-sample Wald CI from a Cox or GLM refit on each frontier member's data. Ignores subgroup-search selection; anti-conservative by construction. -
Split ~ 95% CI– subsample-derived approximation, computed from the empirical SD of averaged half-sample effects across the splits used bycompute_frontier_cis. -
FSBC ~ 95% CI– bias-corrected interval following the bootstrap algorithm of Leon2024fs (eq 7) but treating the selected subgroup as fixed across half-jackknife replicates. The cell showsest (lcl, ucl)whereestis the bias-corrected effect estimate2\hat\beta - \overline{\hat\beta^{(h)}}.
See compute_frontier_cis for the algebra.
Value
Depending on format:
"gt"A
gt_tblobject. ReturnsNULLinvisibly (with a message) if the frontier is unavailable or empty."data.table"A
data.tableof the frontier with effect on the natural scale, anis_selectedlogical column, and effect column renamed to theeffect_measurelabel (e.g.,"HR","OR"). Returns an emptydata.tableif the frontier is unavailable.
See Also
compute_pareto_frontier,
compute_frontier_cis,
plot_pareto_frontier,
sort_subgroups, forestsearch.
Examples
## Not run:
# Survival example
data(gbsg, package = "survival")
fs <- forestsearch(gbsg, ...)
# Basic table (no CIs)
pareto_frontier_table(fs)
pareto_frontier_table(fs, format = "data.table")
# With CIs (compute once, display anywhere)
ci_tab <- compute_frontier_cis(fs, n_splits = 1000, seed = 1)
pareto_frontier_table(fs, ci_table = ci_tab)
plot_pareto_frontier(fs, ci_table = ci_tab) # same intervals
## End(Not run)
Parse sg.harm Factor Names to Expression
Description
Converts ForestSearch factor names (e.g., "er.0", "grade3.1") into human-readable R expressions (e.g., "er <= 0", "grade3 == 1").
Usage
parse_sg_harm_to_expression(sg_harm, fs.est = NULL)
Arguments
sg_harm |
Character vector of factor names from fs.est$sg.harm. |
fs.est |
ForestSearch object (for accessing confs_labels if available). |
Value
Character string expression or NULL if parsing fails.
Plot ForestSearch Results
Description
Dispatches to plot_sg_results for Kaplan-Meier curves,
hazard-ratio forest plots, or combined panels.
Usage
## S3 method for class 'forestsearch'
plot(
x,
type = c("combined", "km", "forest", "summary"),
outcome.name = "Y",
event.name = "Event",
treat.name = "Treat",
...
)
Arguments
x |
A |
type |
Character. Type of plot:
|
outcome.name |
Character. Name of time-to-event column.
Default: |
event.name |
Character. Name of event indicator column.
Default: |
treat.name |
Character. Name of treatment column.
Default: |
... |
Additional arguments passed to |
Value
Invisibly returns the plot result from
plot_sg_results.
See Also
plot_sg_results for full control over appearance,
plot_sg_weighted_km for weighted KM curves,
plot_subgroup_results_forestplot for publication-ready
forest plots.
Examples
## Not run:
fs <- forestsearch(df.analysis = mydata, ...)
# Combined KM + forest plot (default)
plot(fs)
# KM curves only
plot(fs, type = "km")
# Forest plot only
plot(fs, type = "forest")
# With non-standard column names
plot(fs, type = "km",
outcome.name = "os_months",
event.name = "os_event",
treat.name = "treatment")
# With custom labels
plot(fs, sg0_name = "High Risk", sg1_name = "Standard Risk",
treat_labels = c("0" = "Placebo", "1" = "Active Drug"))
## End(Not run)
Plot Method for ForestSearch Forest Plot
Description
Plot Method for ForestSearch Forest Plot
Usage
## S3 method for class 'fs_forestplot'
plot(x, ...)
Arguments
x |
An fs_forestplot object |
... |
Additional arguments (ignored) |
Value
The input x, invisibly.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
fp <- plot_subgroup_results_forestplot(list(fs.est = fs), gbsg,
outcome.name = "rfstime", event.name = "status", treat.name = "hormon")
plot(fp)
## End(Not run)
Plot Method for fs_sg_plot Objects
Description
Plot Method for fs_sg_plot Objects
Usage
## S3 method for class 'fs_sg_plot'
plot(x, which = 1, ...)
Arguments
x |
An fs_sg_plot object |
which |
Character or integer. Which plot to display. Default: 1 (first available) |
... |
Additional arguments passed to plot functions |
Value
x, invisibly. No plot is drawn: a message names the
plot_sg_results() call that produces the requested plot. An error is
raised when a character which does not name an available plot.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
p <- plot_subgroup(fs)
plot(p)
## End(Not run)
Forest plot of an extreme-subgroups summary
Description
Draws one of the vignettes' six forest panels from a
summary.subgroup_sims() object via gg_forest(): the empirical-CI
distribution of the HR point estimate or of the upper 95% bound
UB(HR), for the single-variable panel, the combination + random-
benchmark panel, or the high-risk panel (Pr(UB>=2) filter with the
ITT row anchored first).
Usage
## S3 method for class 'subgroup_sims_summary'
plot(
x,
metric = c("hr", "ub"),
panel = c("single", "combo", "highrisk"),
hr_true = x$hr_true,
base_size = 14,
point_size = 2.2,
line_size = 0.75,
widths = c(3.5, 5, 1.1, 1.2, 1.2, 1.1),
clip_marker = "arrow",
xlim = NULL,
ticks_at = NULL,
xlab = NULL,
footnote = NULL,
...
)
Arguments
x |
A |
metric |
|
panel |
|
hr_true |
True hazard ratio for the reference/vertical line and
footnotes; defaults to the value stored on |
base_size, point_size, line_size |
Passed to |
widths |
Column width vector passed to |
clip_marker |
Passed to |
xlim, ticks_at, xlab, footnote |
|
... |
Further arguments merged over the defaults and passed to
|
Details
Defaults reproduce the committed vignette figures exactly:
metric-specific axis limits, tick positions, reference/vertical line
styling, the four annotation columns (N plus tail probabilities and
the unconditional median), and the panel-specific footnotes. Any
default can be overridden through the named arguments or ...,
which is merged over the assembled argument list before the
gg_forest() call (so e.g. ref_col = "black" or
widths = ... in ... win).
Summaries carrying effect metadata (the generic path of
summary.subgroup_sims(), e.g. from subgroup_glm() fits) take a
scale-aware branch instead: identity-scale measures (MD) plot with
xlog = FALSE and the null line at 0, ratio-scale measures (OR)
with xlog = TRUE and the null line at 1, and both delegate
limits/ticks to gg_forest()'s data-driven defaults (the HR panel
constants can hide binary medians and remain available via
xlim=/ticks_at=); annotation
columns and footnotes are built from the resolved labels and
threshold strings (legacy literals preserved when a pair equals the
legacy pair), and columns for disabled (NA) thresholds are dropped
(the default widths shrink to match; an explicit widths is used
as given). hr_true is interpreted on the fitter's estimate scale
throughout. Legacy summaries produce byte-identical gg_forest()
calls.
Value
The gg_forest() plot object.
See Also
forest_height(), summary.subgroup_sims()
Examples
## Not run:
S <- summary(sims_uniform, hr_true = 0.70)
plot(S, metric = "ub", panel = "combo")
## End(Not run)
Plot Detection Probability Curve
Description
Creates a visualization of the detection probability curve.
Usage
plot_detection_curve(
curve_data,
add_reference_lines = TRUE,
add_threshold_line = TRUE,
title = NULL,
...
)
Arguments
curve_data |
A data.frame from |
add_reference_lines |
Logical. Add horizontal reference lines at 0.05, 0.10, 0.80. Default: TRUE |
add_threshold_line |
Logical. Add vertical line at hr_threshold. Default: TRUE |
title |
Character. Plot title. Default: auto-generated |
... |
Additional arguments passed to plot() |
Value
Invisibly returns the input data.
Examples
## Not run:
curve_data <- generate_detection_curve(n_sg = 60, prop_cens = 0.2)
plot_detection_curve(curve_data)
## End(Not run)
Three-Panel Effect-Distribution Violin from Simulation Results
Description
Builds a three-panel violin plot (ITT / identified \hat H /
identified \hat H^c) of treatment-effect estimates across
simulation replicates produced by
run_simulation_analysis(). Handles both GLM endpoints
(OR, RD, IRR) and survival (HR), optionally inverts effects to a
benefit scale (subgroup_notation = "benefit"), and overlays a
dotted reference line at the DGM-implied true effect in \hat H.
Usage
plot_effect_distribution(
results,
analysis_method = "FS",
effect_measure = c("OR", "HR", "RD", "IRR", "MD"),
subgroup_notation = c("harm", "benefit"),
reference_effect = NA_real_,
title = NULL,
subtitle = NULL,
panel_labels = NULL,
drop_undetected = TRUE,
trim_threshold = 1000,
trim_fraction = 0.01,
ylim = NULL
)
Arguments
results |
|
analysis_method |
Character. Method label in the
|
effect_measure |
Character. One of |
subgroup_notation |
Character. |
reference_effect |
Numeric or |
title |
Character. Plot title. Default: auto-constructed. |
subtitle |
Character or |
panel_labels |
Named character vector with names |
drop_undetected |
Logical. Restrict the |
trim_threshold |
Numeric or |
trim_fraction |
Numeric in (0, 0.5). Fraction of observations
to trim from each tail of each group when trimming triggers.
Default: |
ylim |
Numeric vector of length 2 or |
Details
The MRCT analogue is SGplot_estimates, which works on
the four-panel mrct_region_sims() schema (ITT / training /
testing / subgroup). This function targets the three-panel
run_simulation_analysis() schema that does not involve a
train/test split.
Value
A ggplot2 object. Has attr(p, "panel_data")
set to the long-format data.table used for plotting (for
downstream summaries or diagnostic inspection). When trimming is
active, also has attr(p, "trim_info") containing per-group
diagnostics (n_total, n_trimmed, n_flagged,
raw_mean, raw_sd, trimmed_mean,
trimmed_sd, lower_bound, upper_bound).
See Also
run_simulation_analysis for the simulation
pipeline, SGplot_estimates for the MRCT analogue,
build_estimation_table for the tabular counterpart.
Examples
## Not run:
# Binary endpoint, benefit search, OR scale
true_or_benefit <- 1 / dgm_calibrated$hazard_ratios$harm_subgroup
p <- plot_effect_distribution(
results_alt,
analysis_method = "FS",
effect_measure = "OR",
subgroup_notation = "benefit",
reference_effect = true_or_benefit,
title = "ACTG175 Binary (HTE): OR Estimates Across Simulations",
subtitle = sprintf("n = 1000, 50 sims | truth OR(Q) = %.2f",
true_or_benefit)
)
print(p)
# Survival endpoint, harm search, HR scale (no inversion)
p_surv <- plot_effect_distribution(
results_alt,
effect_measure = "HR",
subgroup_notation = "harm",
reference_effect = dgm_alt$hazard_ratios$harm_subgroup
)
# Clip y-axis to focus on the interesting OR range
p_clipped <- plot_effect_distribution(
results_alt,
effect_measure = "OR",
subgroup_notation = "benefit",
reference_effect = true_or_benefit,
ylim = c(0, 4)
)
## End(Not run)
Covariate involvement across the sweep
Description
Reproduces the supplement's involvement figure: the share of detections whose rule names the anchor, the true partner, and the proxy, against the sweep dimension.
Usage
plot_fs_identification_involvement(x, roles = NULL, palette = NULL)
Arguments
x |
An |
roles |
Character vector. Which covariates to draw; defaults to the anchor/partner/proxy triple. |
palette |
Named character vector of colours, or |
Value
A ggplot object.
Composition of the identified subgroup by structure
Description
Reproduces the supplement's stacked-composition figure.
Usage
plot_fs_identification_structure(x, min_label = 0.03, palette = NULL)
Arguments
x |
An |
min_label |
Numeric. Shares below this are drawn without a label. |
palette |
Named character vector of fills, or |
Value
A ggplot object.
Selected-covariate-pair frequencies
Description
Bar-chart companion to fs_rule_covariate_pairs(), in the same idiom as the
selected-covariate figure: horizontal bars, ordered by share, coloured by
how the signature relates to the true region.
Usage
plot_fs_rule_pairs(
x,
true_covariates = NULL,
top_n = 12L,
detected_only = TRUE
)
Arguments
x |
An |
true_covariates |
Character vector; taken from |
top_n |
Integer. Show only the most frequent |
detected_only |
Passed through. |
Value
A ggplot object.
Plot Kaplan-Meier Survival Difference Bands for ForestSearch Subgroups
Description
Creates Kaplan-Meier survival difference band plots comparing the identified
ForestSearch subgroup (sg.harm) and its complement against the ITT population.
This function wraps plotKM.band_subgroups() from the weightedsurv
package, automatically extracting subgroup definitions from ForestSearch
results.
Usage
plot_km_band_forestsearch(
df,
fs.est = NULL,
sg_cols = NULL,
sg_labels = NULL,
sg_colors = NULL,
itt_color = "azure3",
outcome.name = "tte",
event.name = "event",
treat.name = "treat",
xlabel = "Time",
ylabel = "Survival differences",
yseq_length = 5,
draws_band = 1000,
tau_add = NULL,
by_risk = 6,
risk_cex = 0.75,
risk_delta = 0.035,
risk_pad = 0.015,
ymax_pad = 0.11,
show_legend = TRUE,
legend_pos = "topleft",
legend_cex = 0.75,
ref_subgroups = NULL,
verbose = FALSE
)
Arguments
df |
Data frame. The analysis dataset containing all required variables including subgroup indicator columns. |
fs.est |
A forestsearch object containing the identified subgroup,
or |
sg_cols |
Character vector. Names of columns in |
sg_labels |
Character vector. Subsetting expressions for each subgroup,
corresponding to |
sg_colors |
Character vector. Colors for each subgroup curve,
corresponding to |
itt_color |
Character. Color for ITT population band.
Default: |
outcome.name |
Character. Name of time-to-event column.
Default: |
event.name |
Character. Name of event indicator column.
Default: |
treat.name |
Character. Name of treatment column.
Default: |
xlabel |
Character. X-axis label. Default: |
ylabel |
Character. Y-axis label. Default: |
yseq_length |
Integer. Number of y-axis tick marks.
Default: |
draws_band |
Integer. Number of bootstrap draws for confidence band.
Default: |
tau_add |
Numeric. Time horizon for the plot. If |
by_risk |
Numeric. Interval for risk table. Default: |
risk_cex |
Numeric. Character expansion for risk table text.
Default: |
risk_delta |
Numeric. Vertical spacing for risk table.
Default: |
risk_pad |
Numeric. Padding for risk table. Default: |
ymax_pad |
Numeric. Y-axis maximum padding. Default: |
show_legend |
Logical. Whether to display the legend.
Default: |
legend_pos |
Character. Legend position (e.g., "topleft", "bottomright").
Default: |
legend_cex |
Numeric. Character expansion for legend text.
Default: |
ref_subgroups |
Named list. Optional additional reference subgroups to include. Each element should be a list with:
The function automatically creates indicator columns from the expressions.
Default: |
verbose |
Logical. Print diagnostic messages. Default: |
Details
This function simplifies the workflow of creating KM survival difference band plots for ForestSearch-identified subgroups. It can work in two modes:
Mode 1: With ForestSearch result (fs.est provided)
Extracts the subgroup definition from the ForestSearch result
Creates binary indicator columns (Qrecommend, Brecommend) in
dfGenerates appropriate labels from the subgroup definition
Calls
plotKM.band_subgroups()with configured parameters
Mode 2: With pre-defined columns (sg_cols provided)
Uses existing indicator columns in
dfRequires
sg_labelsandsg_colorsto matchsg_cols
The sg.harm subgroup (Qrecommend) represents patients with questionable
treatment benefit (where treat.recommend == 0 in ForestSearch output).
The complement (Brecommend) represents patients recommended for treatment.
Value
Invisibly returns a list containing:
- df
The modified data frame with subgroup indicators
- sg_cols
Character vector of subgroup column names used
- sg_labels
Character vector of subgroup labels used
- sg_colors
Character vector of colors used
- sg_harm_definition
The subgroup definition extracted from fs.est
- ref_subgroups
The reference subgroups list (if provided)
Subgroup Extraction
When fs.est is provided, the subgroup definition is extracted from:
-
fs.est$grp.consistency$out_sg$sg.harm_label- Human-readable labels -
fs.est$sg.harm- Technical factor names (fallback) -
fs.est$df.est$treat.recommend- Subgroup membership indicator
Note
This function requires the weightedsurv package, which can be
installed from CRAN: install.packages("weightedsurv")
See Also
forestsearch for running the subgroup analysis
plot_sg_weighted_km for weighted KM plots
plot_sg_results for comprehensive subgroup visualization
Examples
## Not run:
# Mode 1: Using ForestSearch result (auto-creates Qrecommend/Brecommend)
# This will use labels "Qrecommend == 1" and "Brecommend == 1"
plot_km_band_forestsearch(
df = df.analysis,
fs.est = fs_result,
outcome.name = "os_months",
event.name = "os_event",
treat.name = "treatment"
)
# Mode 1 with additional reference subgroups (auto-creates columns)
ref_sgs <- list(
age_young = list(subset_expr = "age < 65", color = "brown"),
age_old = list(subset_expr = "age >= 65", color = "grey")
)
plot_km_band_forestsearch(
df = df.analysis,
fs.est = fs_result,
ref_subgroups = ref_sgs,
outcome.name = "os_months",
event.name = "os_event",
treat.name = "treatment",
tau_add = 60,
by_risk = 6
)
# Mode 2: Using pre-defined subgroup columns with expression labels
# Note: sg_labels must be valid R expressions that evaluate against df
df.analysis$age_lt65 <- ifelse(df.analysis$age < 65, 1, 0)
df.analysis$age_ge65 <- ifelse(df.analysis$age >= 65, 1, 0)
df.analysis$Qrecommend <- ifelse(df.analysis$er <= 0, 1, 0)
df.analysis$Brecommend <- ifelse(df.analysis$er > 0, 1, 0)
plot_km_band_forestsearch(
df = df.analysis,
sg_cols = c("age_lt65", "age_ge65", "Qrecommend", "Brecommend"),
sg_labels = c("age < 65", "age >= 65", "er <= 0", "er > 0"),
sg_colors = c("brown", "grey", "blue", "red"),
outcome.name = "os_months",
event.name = "os_event",
treat.name = "treatment",
tau_add = 60,
by_risk = 6
)
## End(Not run)
Combined Pareto-Frontier Plot Across Configurations Sharing a Passing Set
Description
For two or more forestsearch fits whose consistency-passing
candidate sets are identical, produces a single Pareto-frontier
plot annotated with one S<i>: <combo_label> marker per
configuration at each selected subgroup. If two configurations pick
the same subgroup, the markers stack into a single multi-line label
(e.g.\ "S1: hrMaxSG/pareto\nS2: hrMaxSG/both").
Usage
plot_pareto_combined(
fs_list,
combo_labels = NULL,
ci_table_list = NULL,
show_band = TRUE,
xlim = NULL,
tolerance = 1e-06,
verbose = TRUE
)
Arguments
fs_list |
A list of |
combo_labels |
Character vector of length |
ci_table_list |
Optional list of CI tables from
|
show_band |
Logical. Draw the effect-band shading if applicable.
Default |
xlim |
Numeric vector of length 2 or |
tolerance |
Numeric. Per-cell tolerance for the value
equality check ( |
verbose |
Logical. Emit a warning naming the specific
equality-check failure mode when the sets don't match. Default
|
Details
The combined plot is meaningful only when the passing sets match.
When they don't, the function returns NULL with a warning;
the side-by-side view (one panel per configuration) is the
appropriate alternative.
Value
A ggplot object, or NULL (with a warning) if
the passing sets don't satisfy the equality criterion.
Equality criterion
Two passing sets are considered equal when, after sorting by subgroup definition string:
they have the same number of rows;
the same set of subgroup definitions (concatenated
M.*columns), compared as sets (order-independent);within each matched definition, the
hr,N,E, andKvalues agree withintolerance.
Pcons is deliberately excluded from the value check. It can
legitimately drift across rules because the preview sort (which
depends on selection_rule) changes the queue order, which
changes the random-split state each candidate consumes. Drift of
up to ~0.10 between runs on the same subgroup is expected and does
not indicate a real disagreement. Note also that the internal
candidate id m can differ across configurations for the
same reason; m is NOT used for equality.
See Also
compare_selection_rules,
plot_pareto_frontier,
pareto_frontier_table.
Examples
## Not run:
out <- compare_selection_rules(
df.analysis = actg_df,
sg_focus = c("hrMaxSG", "hrMaxSG"),
selection_rule = c("pareto", "both"),
...
)
# Auto-attached by compare_selection_rules():
print(out$plot_combined)
# Or call directly:
p <- plot_pareto_combined(
fs_list = out$fs,
combo_labels = out$combos$label,
ci_table_list = out$ci_tab
)
print(p)
## End(Not run)
Plot the Pareto Frontier of Candidate Subgroups
Description
Produces a 2D scatter of candidate subgroups in (effect, N) space,
with the Pareto frontier drawn as a step polyline and the selected
subgroup highlighted. When ci_table is supplied (the return
value of compute_frontier_cis), horizontal error bars
for the split-derived 95\
Usage
plot_pareto_frontier(
fs,
ci_table = NULL,
show_band = FALSE,
effect_neighborhood = NULL,
label_members = TRUE,
point_size = 3,
xlim = NULL,
xlim_trim = NULL
)
Arguments
fs |
A |
ci_table |
Optional |
show_band |
Logical. If |
effect_neighborhood |
Numeric. Override of the band width;
defaults to the value used in the original |
label_members |
Logical. If |
point_size |
Numeric. Size of frontier-member points.
Default |
xlim |
Optional numeric vector of length 2 controlling the
x-axis range, e.g.\ |
xlim_trim |
Logical or |
Details
The frontier polyline is drawn as a downward step from large-effect
/ small-N candidates to small-effect / large-N candidates. Points
off the frontier (if any are present in the fs$grp.consistency$out_sg$result
table beyond those on the frontier) are not currently plotted; this
function only displays the frontier and its selected member.
For sg_focus = "hrMinSG", the selected subgroup may not lie
on the Pareto frontier (the focus deliberately prefers small
subgroups, which are typically N-dominated). In that case the
selected marker appears off the polyline.
Value
A ggplot object.
See Also
pareto_frontier_table,
compute_frontier_cis, frontier_member_flags.
Examples
## Not run:
p <- plot_pareto_frontier(fs)
print(p)
# With split-derived 95% CIs:
ci_tab <- compute_frontier_cis(fs, n_splits = 1000)
p2 <- plot_pareto_frontier(fs, ci_table = ci_tab)
print(p2)
## End(Not run)
True log-HR psi0(x) along x_pred with region histograms (paper Fig. 3 analogue)
Description
True log-HR psi0(x) along x_pred with region histograms (paper Fig. 3 analogue)
Usage
plot_region_modifier(dgm, file = NULL, main = NULL)
Arguments
dgm |
An MRCT data-generating mechanism returned by
|
file |
Optional character path. When supplied, the plot is written
to this PNG file (900 x 800 px) and the device is closed on exit;
when |
main |
Optional character title for the upper panel. When |
Value
NULL, invisibly. Called for its side effect of drawing a
two-panel plot.
Plot Distribution of Identified Subgroups
Description
Bar chart of subgroups identified across simulations, filtered to
those appearing in at least min_pct of the found simulations.
Supports two column schemas (MRCT and run_simulation_analysis),
top-K capping with a pooled "Other" bar, automatic threshold
halving when no label clears the initial cut, and an n_bars
attribute for adaptive figure-height computation.
Usage
plot_sg_distribution(
results,
min_pct = 5,
title = "Distribution of Identified Subgroups",
wrap_width = 25,
any_col = "any_found",
label_col = "sg_found",
top_k = NULL,
show_other = TRUE,
min_floor_pct = 0.5,
max_halvings = 0L,
bar_label_inside = FALSE,
placeholder_on_empty = FALSE
)
Arguments
results |
|
min_pct |
Numeric. Minimum percentage threshold for display
(0-100). Default: |
title |
Character. Plot title. |
wrap_width |
Integer. Character width for wrapping long
subgroup labels. Default: |
any_col |
Character. Column name of the per-replicate
detection flag. Default: |
label_col |
Character. Column name of the identified-subgroup
label string. Default: |
top_k |
Integer or |
show_other |
Logical. When |
min_floor_pct |
Numeric. Lower bound on automatic threshold
halving. Default: |
max_halvings |
Integer. Maximum number of halvings applied to
|
bar_label_inside |
Logical. Render bar labels inside bars
( |
placeholder_on_empty |
Logical. When no subgroups clear the
threshold (even after halving), return a minimal
"no data" placeholder plot with |
Value
A ggplot2 object, or invisible(NULL) when no
subgroups are found and placeholder_on_empty = FALSE.
When a plot is returned it carries two attributes:
n_barsInteger. Number of bars in the plot (0 in the placeholder case). Use
sgdist_fig_height()to compute an adaptive figure height for knitr.effective_min_pctNumeric. The threshold in effect after any automatic halving.
See Also
sgdist_fig_height for computing an adaptive
figure height from n_bars, and
run_simulation_analysis /
mrct_region_sims for the upstream simulation
engines.
Examples
## Not run:
# MRCT / mrct_region_sims() output (default column schema)
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
plot_sg_distribution(fs$grp.consistency$result_new)
# run_simulation_analysis() output with top-K cap and auto-halving
fs_rows <- results_alt[results_alt$analysis == "FS", ]
p <- plot_sg_distribution(
fs_rows,
any_col = "any.H",
label_col = "sg.def",
top_k = 15L,
show_other = FALSE,
max_halvings = 2L,
bar_label_inside = TRUE,
placeholder_on_empty = TRUE,
title = "HTE: Subgroups Identified (H-hat)"
)
fig_h <- sgdist_fig_height(attr(p, "n_bars"))
## End(Not run)
Plot Forest Plot of Hazard Ratios
Description
Creates a forest plot showing hazard ratios with confidence intervals.
Usage
plot_sg_forest(hr_estimates, sg0_name, sg1_name, colors, title = NULL, ...)
Arguments
hr_estimates |
Data frame with HR estimates |
sg0_name |
Character. Label for H subgroup |
sg1_name |
Character. Label for Hc subgroup |
colors |
List. Color specifications |
title |
Character. Plot title |
... |
Additional arguments |
Value
Invisible NULL (creates plot as side effect)
Plot GLM Subgroup Outcomes
Description
Creates a publication-ready visualization of identified subgroup outcomes
for binary or continuous endpoints. This is the GLM counterpart to
plot_sg_weighted_km for survival endpoints.
Usage
plot_sg_glm_outcomes(
fs.est,
fs_bc = NULL,
outcome.name,
treat.name,
offset.name = NULL,
effect_measure = NULL,
outcome_type = NULL,
adjust_covariates = NULL,
sg0_name = "Questionable",
sg1_name = "Recommend",
E.name = "Treatment",
C.name = "Control",
bar_colors = c("steelblue", "coral"),
show_effect = TRUE,
show_bc = TRUE,
show_itt = TRUE,
conf.level = 0.95,
title = NULL,
subtitle = NULL,
verbose = FALSE
)
## S3 method for class 'fs_binary_plot'
print(x, ...)
## S3 method for class 'fs_binary_plot'
plot(x, ...)
Arguments
fs.est |
A forestsearch object containing |
fs_bc |
Optional. Bootstrap results from
|
outcome.name |
Character. Name of outcome column. |
treat.name |
Character. Name of treatment column. |
offset.name |
Character or |
effect_measure |
Character or |
outcome_type |
Character or |
adjust_covariates |
Character vector or |
sg0_name |
Character. Label for H subgroup (treat.recommend == 0).
Default: |
sg1_name |
Character. Label for Hc subgroup (treat.recommend == 1).
Default: |
E.name |
Character. Label for experimental/treatment arm.
Default: |
C.name |
Character. Label for control arm.
Default: |
bar_colors |
Character vector of length 2. Colors for control and
treatment bars. Default: |
show_effect |
Logical. Annotate with effect estimate. Default
|
show_bc |
Logical. Annotate with bias-corrected estimate when
|
show_itt |
Logical. Include an ITT (full population) panel.
Default |
conf.level |
Numeric. Confidence level for intervals.
Default |
title |
Character or |
subtitle |
Character or |
verbose |
Logical. Print diagnostic messages. Default |
x |
An fs_binary_plot object. |
... |
Additional arguments (unused). |
Details
For binary outcomes, displays grouped bar charts of event rates by treatment arm within each subgroup. For continuous outcomes, displays point-and-errorbar charts of means (+/- SE) by treatment arm. Both are annotated with the effect estimate and optional bias-corrected estimate from the bootstrap.
Value
A list of class "fs_binary_plot" containing:
- plot
A ggplot2 object.
- data
Data frame of per-arm summaries used for the plot.
- effects
Data frame of effect estimates per subgroup.
- sg_definition
Character. Subgroup definition label.
- figure_note
Character. Auto-generated figure note.
Examples
## Not run:
fs <- forestsearch(df, outcome_type = "binary", effect_measure = "OR",
...)
result <- plot_sg_glm_outcomes(fs, outcome.name = "status",
treat.name = "hormon")
print(result)
## End(Not run)
Plot Kaplan-Meier Survival Curves for Subgroups
Description
Creates side-by-side Kaplan-Meier survival curves for the H and Hc subgroups.
Usage
plot_sg_km(
df_H,
df_Hc,
outcome.name,
event.name,
treat.name,
by.risk,
sg0_name,
sg1_name,
treat_labels,
colors,
show_ci = TRUE,
show_logrank = TRUE,
show_hr = TRUE,
hr_estimates = NULL,
conf.level = 0.95,
title = NULL,
...
)
Arguments
df_H |
Data frame for H subgroup |
df_Hc |
Data frame for Hc subgroup |
outcome.name |
Character. Outcome variable name |
event.name |
Character. Event indicator name |
treat.name |
Character. Treatment variable name |
by.risk |
Numeric. Risk table interval |
sg0_name |
Character. Label for H subgroup |
sg1_name |
Character. Label for Hc subgroup |
treat_labels |
Named character vector. Treatment labels |
colors |
List. Color specifications |
show_ci |
Logical. Show confidence intervals |
show_logrank |
Logical. Show log-rank p-value |
show_hr |
Logical. Show HR annotation |
hr_estimates |
Data frame with HR estimates |
conf.level |
Numeric. Confidence level |
title |
Character. Plot title |
... |
Additional arguments |
Value
Invisible NULL (creates plot as side effect)
Plot ForestSearch Subgroup Results
Description
Creates comprehensive visualizations of subgroup results from ForestSearch,
including Kaplan-Meier survival curves, hazard ratio comparisons, and
summary statistics. This function is designed to work with the output
from forestsearch, specifically the df.est component.
Usage
plot_sg_results(
fs.est,
outcome.name = "Y",
event.name = "Event",
treat.name = "Treat",
plot_type = c("combined", "km", "forest", "summary"),
by.risk = NULL,
conf.level = 0.95,
est.scale = c("hr", "1/hr"),
sg0_name = "Questionable (H)",
sg1_name = "Recommend (H^c)",
treat_labels = c(`0` = "Control", `1` = "Treatment"),
colors = NULL,
title = NULL,
show_events = TRUE,
show_ci = TRUE,
show_logrank = TRUE,
show_hr = TRUE,
verbose = FALSE,
...
)
Arguments
fs.est |
A forestsearch object or list containing at minimum:
|
outcome.name |
Character. Name of time-to-event outcome column. Default: "Y" |
event.name |
Character. Name of event indicator column (1=event, 0=censored). Default: "Event" |
treat.name |
Character. Name of treatment column (1=treatment, 0=control). Default: "Treat" |
plot_type |
Character. Type of plot to create. One of:
|
by.risk |
Numeric. Risk interval for KM survival curves. Default: NULL (auto-calculated) |
conf.level |
Numeric. Confidence level for intervals. Default: 0.95 |
est.scale |
Character. Effect scale: "hr" (hazard ratio) or "1/hr" (inverse). Default: "hr" |
sg0_name |
Character. Label for subgroup 0 (harm/questionable). Default: "Questionable (H)" |
sg1_name |
Character. Label for subgroup 1 (recommend/complement). Default: "Recommend (H^c)" |
treat_labels |
Named character vector. Labels for treatment arms. Default: c("0" = "Control", "1" = "Treatment") |
colors |
Named character vector. Colors for plot elements. Default: uses package defaults |
title |
Character. Main plot title. Default: auto-generated |
show_events |
Logical. Show event counts on KM curves. Default: TRUE |
show_ci |
Logical. Show confidence intervals. Default: TRUE |
show_logrank |
Logical. Show log-rank p-value. Default: TRUE |
show_hr |
Logical. Show hazard ratio annotation. Default: TRUE |
verbose |
Logical. Print diagnostic messages. Default: FALSE |
... |
Additional arguments passed to plotting functions. |
Details
The function extracts subgroup membership from fs.est$df.est$treat.recommend:
-
treat.recommend == 0: Harm/questionable subgroup (H) -
treat.recommend == 1: Recommend/complement subgroup (H^c)
For est.scale = "1/hr", treatment labels and subgroup interpretation
are reversed to maintain clinical interpretability.
Value
An object of class fs_sg_plot containing:
- plots
List of ggplot2 or base R plot objects
- summary
Data frame of subgroup summary statistics
- hr_estimates
Data frame of hazard ratio estimates
- call
The matched call
Kaplan-Meier Plots
When plot_type = "km", creates side-by-side survival curves for:
The identified subgroup (H) with treatment vs control
The complement subgroup (H^c) with treatment vs control
Forest Plot
When plot_type = "forest", creates a forest plot showing hazard
ratios with confidence intervals for: ITT population, H subgroup,
and H^c complement.
See Also
forestsearch for running the subgroup analysis
sg_consistency_out for consistency evaluation
plot_subgroup_results_forestplot for publication-ready forest plots
Examples
## Not run:
# Run ForestSearch analysis
fs <- forestsearch(
df.analysis = my_data,
outcome.name = "os_time",
event.name = "os_event",
treat.name = "treatment",
confounders.name = c("age_cat", "stage", "biomarker")
)
# Create combined visualization
result <- plot_sg_results(
fs.est = fs,
outcome.name = "os_time",
event.name = "os_event",
treat.name = "treatment"
)
# View the Kaplan-Meier plots only
plot_sg_results(fs, plot_type = "km")
# Customize labels
plot_sg_results(
fs,
sg0_name = "High Risk",
sg1_name = "Standard Risk",
treat_labels = c("0" = "Placebo", "1" = "Active Drug")
)
## End(Not run)
Plot Summary Statistics Panel
Description
Creates a summary panel with subgroup characteristics.
Usage
plot_sg_summary_panel(
summary_stats,
hr_estimates,
sg0_name,
sg1_name,
colors,
...
)
Arguments
summary_stats |
Data frame with summary statistics |
hr_estimates |
Data frame with HR estimates |
sg0_name |
Character. Label for H subgroup |
sg1_name |
Character. Label for Hc subgroup |
colors |
List. Color specifications |
... |
Additional arguments |
Value
Invisible NULL (creates plot as side effect)
Plot Weighted Kaplan-Meier Curves for ForestSearch Subgroups
Description
Creates weighted Kaplan-Meier survival curves for the identified subgroups
(H and Hc) using the weightedsurv package, matching the pattern used in
sg_consistency_out().
Usage
plot_sg_weighted_km(
fs.est,
fs_bc = NULL,
outcome.name = "Y",
event.name = "Event",
treat.name = "Treat",
by.risk = NULL,
sg0_name = NULL,
sg1_name = NULL,
conf.int = TRUE,
show.logrank = TRUE,
show.cox = TRUE,
show.cox.bc = TRUE,
put.legend.lr = "topleft",
ymax = 1.05,
xmed.fraction = 0.65,
hr_bc_position = "bottomright",
hr_bc_cex = 0.725,
title = NULL,
verbose = FALSE
)
Arguments
fs.est |
A forestsearch object containing |
fs_bc |
Optional. Bootstrap results from |
outcome.name |
Character. Name of time-to-event column.
Default: |
event.name |
Character. Name of event indicator column.
Default: |
treat.name |
Character. Name of treatment column.
Default: |
by.risk |
Numeric. Risk interval for plotting. Default: |
sg0_name |
Character. Label for H subgroup (treat.recommend == 0).
Default: |
sg1_name |
Character. Label for Hc subgroup (treat.recommend == 1).
Default: |
conf.int |
Logical. Show confidence intervals. Default: |
show.logrank |
Logical. Show log-rank test. Default: |
show.cox |
Logical. Show unadjusted Cox HR from weightedsurv.
Default: |
show.cox.bc |
Logical. Show bootstrap bias-corrected HR annotation
(requires |
put.legend.lr |
Character. Legend position. Default: "topleft" |
ymax |
Numeric. Max y-axis value. Default: 1.05 |
xmed.fraction |
Numeric. Fraction for median lines. Default: 0.65 |
hr_bc_position |
Character. Position for bias-corrected HR annotation. One of "bottomright", "bottomleft", "topright", "topleft". Default: "bottomright" |
hr_bc_cex |
Numeric. Character expansion factor for bias-corrected HR annotation text. Default: 0.725 (matches weightedsurv cox.cex default) |
title |
Character. Overall plot title. Default: |
verbose |
Logical. Print diagnostic messages. Default: |
Details
This function uses the exact same calling pattern as plot_subgroup()
in the ForestSearch package. Column names are mapped internally to the
standard names (Y, Event, Treat) expected by weightedsurv.
Subgroup definitions are automatically extracted from the forestsearch object if available:
-
fs$grp.consistency$out_sg$sg.harm_label- Human-readable labels -
fs$sg.harm- Technical factor names (fallback)
HR display options controlled by show.cox and show.cox.bc:
Both TRUE (default): Shows unadjusted HR from weightedsurv AND bias-corrected HR annotation
-
show.cox = TRUE, show.cox.bc = FALSE: Shows only unadjusted HR -
show.cox = FALSE, show.cox.bc = TRUE: Shows only bias-corrected HR Both FALSE: Shows neither HR estimate
Value
Invisibly returns a list with subgroup data frames and counting data
Examples
## Not run:
# After running forestsearch - auto-extracts subgroup definition
plot_sg_weighted_km(fs.est = fs)
# With bootstrap bias-corrected estimates
plot_sg_weighted_km(fs.est = fs, fs_bc = fs_bootstrap)
# With custom column names
plot_sg_weighted_km(
fs.est = fs,
outcome.name = "time_months",
event.name = "status",
treat.name = "hormon"
)
## End(Not run)
Plot Spline Treatment Effect Function
Description
Plot Spline Treatment Effect Function
Usage
plot_spline_treatment_effect(dgm_result, add_points = TRUE)
Arguments
dgm_result |
Result object from generate_aft_dgm_flex with spline |
add_points |
Logical; add observed data points. Default TRUE |
Value
Called for its side effect of producing a base-R plot. Returns
NULL invisibly.
Examples
## Not run:
library(survival)
df <- survival::gbsg
dgm <- generate_aft_dgm_flex(df, outcome.name = "rfstime",
event.name = "status", treat.name = "hormon",
confounders.name = c("age", "meno", "nodes"))
plot_spline_treatment_effect(dgm)
## End(Not run)
Plot Subgroup Survival Curves
Description
Plots weighted Kaplan-Meier survival curves for a specified subgroup and its complement using the weightedsurv package.
Usage
plot_subgroup(df.sub, df.subC, by.risk, confs_labels, this.1_label, top_result)
Arguments
df.sub |
A data frame containing data for the subgroup of interest. |
df.subC |
A data frame containing data for the complement subgroup. |
by.risk |
Numeric. The risk interval for plotting (passed to |
confs_labels |
Named character vector. Covariate label mapping (not used directly in this function, but may be used for labeling). |
this.1_label |
Character. Label for the subgroup being plotted. |
top_result |
Data frame row. The top subgroup result row, expected to contain a |
Plot Subgroup Analysis Results
Description
Creates diagnostic plots for subgroup treatment effects from df_super object
Usage
plot_subgroup_effects(
df_super,
z,
hrz_crit = 0,
log.hrs = NULL,
ahr_empirical = NULL,
plot_type = c("both", "profile", "ahr"),
add_rug = TRUE,
zpoints_by = 1,
...
)
Arguments
df_super |
A data frame containing subgroup analysis results with columns: loghr_po (log hazard ratios), and optionally theta_1 and theta_0 (treatment-specific parameters) |
z |
Character string specifying the column name to use as the subgroup score (e.g., "z_age", "z_size", "subgroup"). Required. |
hrz_crit |
Critical hazard ratio threshold for defining optimal subgroup. Default is 1 (HR=1 on log scale is 0). |
log.hrs |
Optional vector of reference log hazard ratios to display as horizontal lines. Default is NULL. |
ahr_empirical |
Optional empirical average hazard ratio to display. If NULL, calculated from data. Default is NULL. |
plot_type |
Character string specifying plot type: "both" (default), "profile", or "ahr". |
add_rug |
Logical indicating whether to add rug plot of z values. Default is TRUE. |
zpoints_by |
Step size for z-axis grid when calculating AHR curves. Default is 1. |
... |
Additional graphical parameters passed to plot() |
Details
The function creates up to two plots:
Treatment effect profile: Shows log hazard ratio as function of z
Average hazard ratio curve: Shows AHR for subgroups z >= threshold
The "optimal" subgroup is defined as patients with z >= cut.zero, where cut.zero is the minimum z value with favorable treatment effect (loghr < hrz_crit).
Value
A list containing:
cut.zero |
The minimum z value where loghr_po < hrz_crit |
AHR_opt |
Average hazard ratio for optimal subgroup (z >= cut.zero) |
zpoints |
Grid of z values used for AHR calculations |
HR.zpoints |
AHR for population with z >= zpoints |
HRminus.zpoints |
AHR for population with z <= zpoints |
HR2.zpoints |
Alternative AHR calculation for z >= zpoints |
HRminus2.zpoints |
Alternative AHR calculation for z <= zpoints |
Examples
## Not run:
# Using z_age as the subgroup score
results <- plot_subgroup_effects(dgm_spline$df_super, z = "z_age", hrz_crit = 0)
# Using subgroup identifier
results <- plot_subgroup_effects(dgm_spline$df_super, z = "subgroup", hrz_crit = 0)
# With reference lines
results <- plot_subgroup_effects(dgm_spline$df_super, z = "z_size",
hrz_crit = 0,
log.hrs = c(-0.5, 0, 0.5))
# Only AHR plot
results <- plot_subgroup_effects(dgm_spline$df_super, z = "z_pgr",
plot_type = "ahr")
## End(Not run)
Plot Subgroup Results Forest Plot
Description
Generates a comprehensive forest plot showing:
ITT (Intent-to-Treat) population estimate
Reference subgroups (e.g., by biomarker levels)
Post-hoc identified subgroups with bias-corrected estimates
Cross-validation agreement metrics as annotations
Usage
plot_subgroup_results_forestplot(
fs_results,
df_analysis,
subgroup_list = NULL,
outcome.name,
event.name,
treat.name,
E.name = "Experimental",
C.name = "Control",
est.scale = "hr",
xlog = TRUE,
title_text = NULL,
arrow_text = c("Favors Experimental", "Favors Control"),
footnote_text = c("Eg 80% of training found SG: 70% of B (+) also B in CV testing"),
xlim = c(0.25, 1.5),
ticks_at = c(0.25, 0.7, 1, 1.5),
show_cv_metrics = TRUE,
cv_source = c("auto", "kfold", "oob", "both"),
posthoc_colors = c("powderblue", "beige"),
reference_colors = c("yellow", "powderblue"),
ci_column_spaces = 20,
conf.level = 0.95,
theme = NULL,
outcome_type = NULL,
effect_measure = NULL,
offset.name = NULL,
adjust_covariates = NULL,
extreme_ci_cap = 1.5,
xlim_method = c("clinical", "data")
)
Arguments
fs_results |
List. A list containing ForestSearch analysis results with elements:
|
df_analysis |
Data frame. The analysis dataset with outcome, event, and treatment variables. |
subgroup_list |
List. Named list of subgroup definitions to include in the plot. Each element should be a list with:
|
outcome.name |
Character. Name of the survival time variable. |
event.name |
Character. Name of the event indicator variable. |
treat.name |
Character. Name of the treatment variable. |
E.name |
Character. Label for experimental arm (default: "Experimental"). |
C.name |
Character. Label for control arm (default: "Control"). |
est.scale |
Character. Estimate scale: "hr" or "1/hr" (default: "hr"). |
xlog |
Logical. If TRUE (default), the x-axis is plotted on a logarithmic scale. This is standard for hazard ratio forest plots where equal distances represent equal relative effects. |
title_text |
Character. Plot title (default: NULL). |
arrow_text |
Character vector of length 2. Arrow labels for forest plot (default: c("Favors Experimental", "Favors Control")). |
footnote_text |
Character vector. Footnote text for the plot explaining CV metrics (default provides CV interpretation guidance; set to NULL to omit). |
xlim |
Numeric vector of length 2. X-axis limits (default: c(0.25, 1.5)). |
ticks_at |
Numeric vector. X-axis tick positions (default: c(0.25, 0.70, 1.0, 1.5)). |
show_cv_metrics |
Logical. Whether to show cross-validation metrics (default: TRUE if fs_kfold or fs_OOB available). |
cv_source |
Character. Source for CV metrics: "auto" (default, uses both if available, otherwise whichever is present), "kfold" (use fs_kfold only), "oob" (use fs_OOB only), or "both" (explicitly use both fs_kfold and fs_OOB, with K-fold first then OOB). |
posthoc_colors |
Character vector. Colors for post-hoc subgroup rows (default: c("powderblue", "beige")). |
reference_colors |
Character vector. Colors for reference subgroup rows (default: c("yellow", "powderblue")). |
ci_column_spaces |
Integer. Number of spaces for the CI plot column width. More spaces = wider CI column (default: 20). |
conf.level |
Numeric. Confidence level for intervals (default: 0.95 for 95% CI). Used to calculate the z-multiplier as qnorm(1 - (1 - conf.level)/2). |
theme |
An fs_forest_theme object from |
outcome_type |
Character or |
effect_measure |
Character or |
offset.name |
Character or |
adjust_covariates |
Character vector or |
extreme_ci_cap |
Numeric. Multiplier for outlier-resistant axis
limits when |
xlim_method |
Character. Method for computing default axis limits
for GLM measures when |
Details
Creates a publication-ready forest plot displaying identified subgroups with hazard ratios, bias-corrected estimates, and cross-validation metrics. This wrapper integrates ForestSearch results with the forestploter package.
ForestSearch Labeling Convention
ForestSearch identifies subgroups based on hazard ratio thresholds:
-
sg.harm: Contains the definition of the "harm" or "questionable" subgroup (H) -
treat.recommend == 0: Patient is IN the harm subgroup (H) -
treat.recommend == 1: Patient is in the COMPLEMENT subgroup (Hc, typically benefit)
For est.scale = "hr" (searching for harm):
H (treat.recommend=0): Subgroup defined by sg.harm with elevated HR (harm/questionable)
Hc (treat.recommend=1): Complement of sg.harm (potential benefit)
For est.scale = "1/hr" (searching for benefit):
Roles are reversed: H becomes the benefit group
Value
A list containing:
- plot
The forestploter grob object (can be rendered with plot())
- data
The data frame used for the forest plot
- row_types
Character vector of row types for styling reference
- cv_metrics
Cross-validation metrics text (if available)
See Also
forestsearch for running the subgroup analysis
forestsearch_bootstrap_dofuture for bootstrap bias correction
forestsearch_Kfold for cross-validation
create_forest_theme for customizing plot appearance
render_forestplot for rendering the plot
Examples
## Not run:
# Load ForestSearch results
load("fs_analysis_results.Rdata") # Contains fs.est, fs_bc, fs_kfold
# Define subgroups to display
subgroups <- list(
bm_gt1 = list(
subset_expr = "BM > 1",
name = "BM > 1",
type = "reference"
),
bm_gt1_size_gt19 = list(
subset_expr = "BM > 1 & tmrsize > 19",
name = "BM > 1 & Tumor Size > 19",
type = "posthoc"
)
)
# Create the forest plot with default theme
result <- plot_subgroup_results_forestplot(
fs_results = list(fs.est = fs.est, fs_bc = fs_bc, fs_kfold = fs_kfold),
df_analysis = df_itt,
subgroup_list = subgroups,
outcome.name = "os_time",
event.name = "os_event",
treat.name = "combo",
E.name = "Experimental+CT",
C.name = "CT"
)
# Create with custom theme for larger plot
large_theme <- create_forest_theme(
base_size = 14,
row_padding = c(6, 4),
cv_fontsize = 11
)
result <- plot_subgroup_results_forestplot(
fs_results = list(fs.est = fs.est, fs_bc = fs_bc, fs_kfold = fs_kfold),
df_analysis = df_itt,
subgroup_list = subgroups,
outcome.name = "os_time",
event.name = "os_event",
treat.name = "combo",
theme = large_theme
)
# Display the plot
render_forestplot(result)
## End(Not run)
Predicted DINA values for new covariate data.
Description
Returns the estimated treatment-effect contrast
\hat\tau(x) = \hat\beta_0 + x^\top \hat\beta_{1:d}, on the natural
parameter scale of the fitted family (mean difference for Gaussian, log
odds ratio for binomial, log mean ratio for Poisson, log hazard ratio for
Cox).
Usage
## S3 method for class 'dina'
predict(object, newdata, ...)
Arguments
object |
an object of class |
newdata |
a numeric matrix or data frame with the same number of
columns as the training |
... |
unused. |
Value
numeric vector of predicted treatment-effect contrasts, one per
row of newdata.
Predict Corrected Null FPR
Description
Uses a calibrated L_{\text{eff}} model to predict the
procedure-level FPR at arbitrary sample sizes and thresholds.
Usage
predict_fpr_corrected(
calibration,
theta,
d_eff,
c1,
c2,
N,
effect_scale = "ratio"
)
Arguments
calibration |
An |
theta |
Numeric. True treatment effect (1.0 for null). |
d_eff |
Numeric. Effective information in the subgroup. |
c1 |
Numeric. Screening threshold. |
c2 |
Numeric. Consistency threshold. |
N |
Integer. Total sample size. |
effect_scale |
Character. |
Value
A list with components:
- P1
Per-subgroup detection probability.
- L_eff
Effective number of candidates at this N.
- fpr_corrected
Corrected procedure-level FPR.
Examples
cal <- calibrate_L_eff(
N = c(200, 500, 700, 1000),
P1 = c(0.152, 0.109, 0.091, 0.071),
sim_fpr = c(0.185, 0.220, 0.270, 0.280)
)
# Predict FPR at N = 800
d <- d_eff_binary(n_sg = round(0.30 * 800), p_event = 0.30)
predict_fpr_corrected(cal, theta = 1.0, d_eff = d,
c1 = 1.25, c2 = 1.25, N = 800)
Prepare Censoring Model Parameters
Description
Constructs the censoring model object and appends per-subject counterfactual
censoring linear predictors (lin_pred_cens_0, lin_pred_cens_1)
to the super-population data frame.
Usage
prepare_censoring_model(
df_work,
cens_type,
cens_params,
df_super,
select_censoring = TRUE,
cens_intercept_only = FALSE,
verbose = TRUE
)
Arguments
df_work |
Working data frame (output of |
cens_type |
Character. |
cens_params |
Named list of user-supplied censoring parameters. |
df_super |
Super-population data frame; receives
|
select_censoring |
Logical. Selects among three censoring modes:
|
cens_intercept_only |
Logical. Only honored in force-fit mode
(
Setting |
verbose |
Logical. If |
Details
Linear predictor convention
lin_pred_cens_0 and lin_pred_cens_1 store the
covariate contribution only — i.e. \gamma_c' X, with the
intercept \mu_c excluded. This matches the convention used for the
outcome model (lin_pred_0, lin_pred_1 = \gamma' X,
no intercept) computed in calculate_linear_predictors().
simulate_from_dgm() reconstructs the full log-censoring time as:
\log C = \mu_c + \delta + \tau_c \epsilon + \gamma_c' X
where \mu_c = params$censoring$mu,
\delta = cens_adjust,
\tau_c = params$censoring$tau, and
\gamma_c' X = lin_pred_cens_{0|1}.
When select_censoring = TRUE, predict(survreg, type = "linear")
returns the full linear predictor \mu_c + \gamma_c' X. The stored
intercept \mu_c is therefore subtracted before writing
lin_pred_cens_*, so that simulate_from_dgm() can add
params$censoring$mu exactly once. Omitting this subtraction causes
\mu_c to be counted twice, producing astronomically large censoring
times and universal censoring.
When select_censoring = FALSE with a Weibull/lognormal
cens_type, the intercept-only model has zero covariate contribution,
so lin_pred_cens_0 = lin_pred_cens_1 = 0. Storing mu instead
of 0 causes the same double-counting.
Value
A named list:
- cens_model
List of censoring distribution parameters stored in
dgm$model_params$censoring.- df_super
Updated super-population data frame with
lin_pred_cens_0andlin_pred_cens_1appended. These hold covariate contributions only (\gamma_c' X); the intercept is excluded.
Prepare Censoring Model Parameters
Description
Prepare Censoring Model Parameters
Usage
prepare_censoring_model_legacy(df_work, cens_type, cens_params, df_super)
Prepare Data for Subgroup Search
Description
Cleans data by removing missing values and extracting components
Usage
prepare_search_data(Y, Event, Treat, Z)
Prepare subgroup data for analysis
Description
Splits a data frame into two subgroups based on a flag and treatment scale.
Usage
prepare_subgroup_data(df, SG_flag, est.scale, treat.name)
Arguments
df |
Data frame. |
SG_flag |
Character. Name of subgroup flag variable. |
est.scale |
Character. Effect scale ("hr" or "1/hr"). |
treat.name |
Character. Name of treatment variable. |
Value
List with subgroup data frames and treatment variable name.
Examples
df <- data.frame(
treat = c(0, 1, 0, 1, 0),
sg_flag = c(1, 1, 0, 0, 1)
)
result <- prepare_subgroup_data(df, SG_flag = "sg_flag",
est.scale = "hr", treat.name = "treat")
nrow(result$df_1)
Prepare Working Dataset with Processed Covariates
Description
Prepare Working Dataset with Processed Covariates
Usage
prepare_working_dataset(
data,
outcome_var,
event_var,
treatment_var,
continuous_vars,
factor_vars,
standardize,
continuous_vars_cens,
factor_vars_cens,
verbose
)
Print method for cox_ahr_cde objects
Description
Print method for cox_ahr_cde objects
Usage
## S3 method for class 'cox_ahr_cde'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (not used). |
Value
Invisibly returns the input object.
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
res <- cox_ahr_cde_analysis(df,
outcome.name = "rfstime", event.name = "status", treat.name = "hormon",
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"))
print(res)
## End(Not run)
Print a bagged DINA fit.
Description
Extends print.dina() with bag count and per-bag cross-fitting info.
Usage
## S3 method for class 'dina_bagged'
print(x, digits = max(3L, getOption("digits") - 3L), ...)
Arguments
x |
an object of class |
digits |
number of digits to show for coefficients. |
... |
unused. |
Value
invisibly returns x.
Print a dina_subgroup result.
Description
Print a dina_subgroup result.
Usage
## S3 method for class 'dina_subgroup'
print(x, digits = max(3L, getOption("digits") - 3L), ...)
Arguments
x |
a |
digits |
number of digits for numeric summary. |
... |
unused. |
Value
invisibly returns x.
Print a dina_subgroup_bootstrap result.
Description
Print a dina_subgroup_bootstrap result.
Usage
## S3 method for class 'dina_subgroup_bootstrap'
print(x, digits = max(3L, getOption("digits") - 3L), max_select = 5L, ...)
Arguments
x |
a |
digits |
number of digits for numeric summary. |
max_select |
maximum number of rows of the selection-frequency
table to display. Default |
... |
unused. |
Value
invisibly returns x.
Print a dina_subgroup_refit result.
Description
Print a dina_subgroup_refit result.
Usage
## S3 method for class 'dina_subgroup_refit'
print(x, digits = max(3L, getOption("digits") - 3L), ...)
Arguments
x |
a |
digits |
number of digits for the numeric summary. |
... |
unused. |
Value
invisibly returns x.
Print Method for forestsearch Objects
Description
Displays a concise summary of ForestSearch results including the identified subgroup definition, consistency metrics, algorithm details, and computation time.
Usage
## S3 method for class 'forestsearch'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (currently unused). |
Value
Invisibly returns x.
Post-selection inference
When the fit carries multiplier-resampling results (mr_inference = TRUE,
attached as x$mr_inference) and the field block ran
(ci_method = "field", the default), a further block reports: the
one-sided 95% lower bound for the identified subgroup (the field); the
one-sided 95% upper bound for its complement (field-s, the studentized
complement field); the Bonferroni pair for a joint two-subgroup
statement; the infinitesimal-jackknife (IJ) two-term two-sided interval,
as a secondary summary; and the re-selection frequency
\hat p(\widehat H) as a descriptive diagnostic. Bounds are read by
their location relative to a clinically meaningful effect size, not as
tests at the null. With ci_method = "ij" only the two-sided interval is
shown; without MR results the output is unchanged. The constructions and
their simulation operating characteristics are described in León and
Anderson (2026), arXiv:2609.38361.
See Also
summary.forestsearch for detailed output,
plot.forestsearch for visualization.
Examples
## Not run:
fs <- forestsearch(df.analysis = mydata, ...)
print(fs)
# or simply:
fs
## End(Not run)
Print Method for forestsearch_comparison
Description
Short summary of the comparison. Use print(x$plot_grid) for
the side-by-side plot and cat(x$console[[i]]) for each
combo's diagnostic output.
Usage
## S3 method for class 'forestsearch_comparison'
print(x, ...)
Arguments
x |
A |
... |
Ignored. |
Value
x, invisibly. Called for its side effect of printing the
comparison summary.
Print method for fpr_calibration objects
Description
Print method for fpr_calibration objects
Usage
## S3 method for class 'fpr_calibration'
print(x, ...)
Arguments
x |
An |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print Method for ForestSearch Bootstrap Results
Description
Prints a concise but informative summary of an fs_bootstrap
object returned by forestsearch_bootstrap_dofuture.
Covers identification rate, agreement with the primary subgroup,
top identified subgroups, and bias-corrected effect estimates.
Usage
## S3 method for class 'fs_bootstrap'
print(x, top_n = 3L, ...)
Arguments
x |
An |
top_n |
Integer. Maximum number of top identified subgroups to print. Default 3. |
... |
Additional arguments (ignored; present for S3 consistency). |
Details
The print method is deliberately richer than the corresponding
print.fs_kfold / print.fs_tenfold methods because the
bootstrap is the primary inferential machinery of the package, while
the CV routines are complementary diagnostics. For the full
per-factor breakdown, consistency distribution, size distribution,
and GRF cut tabulation, call summarize_bootstrap_subgroups.
Value
Invisibly returns x.
See Also
forestsearch_bootstrap_dofuture to produce the object.
summarize_bootstrap_subgroups for full tabulations.
Examples
## Not run:
fs <- forestsearch(
df.analysis = mydata,
confounders.name = c("age", "sex", "biomarker"),
outcome.name = "time", event.name = "status", treat.name = "treat"
)
fs_bc <- forestsearch_bootstrap_dofuture(fs, nb_boots = 500)
print(fs_bc) # rich default summary
print(fs_bc, top_n = 5)
## End(Not run)
Print Method for fs_cv_summary Objects
Description
Print Method for fs_cv_summary Objects
Usage
## S3 method for class 'fs_cv_summary'
print(x, n = 5L, ...)
Arguments
x |
An |
n |
Integer. Maximum rows to print from each raw data frame
(default: |
... |
Additional arguments (currently unused). |
Value
Invisibly returns x.
Print method for fs_dgm_scale objects
Description
Print method for fs_dgm_scale objects
Usage
## S3 method for class 'fs_dgm_scale'
print(x, ...)
Arguments
x |
An object of class |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print method for fs_fdr_report objects
Description
Print method for fs_fdr_report objects
Usage
## S3 method for class 'fs_fdr_report'
print(x, ...)
Arguments
x |
An |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print Method for ForestSearch Forest Theme
Description
Print Method for ForestSearch Forest Theme
Usage
## S3 method for class 'fs_forest_theme'
print(x, ...)
Arguments
x |
An fs_forest_theme object |
... |
Additional arguments (ignored) |
Value
The input x, invisibly.
Examples
theme <- create_forest_theme()
print(theme)
Print Method for ForestSearch Forest Plot
Description
Print Method for ForestSearch Forest Plot
Usage
## S3 method for class 'fs_forestplot'
print(x, ...)
Arguments
x |
An fs_forestplot object |
... |
Additional arguments (ignored) |
Value
The input x, invisibly.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
fp <- plot_subgroup_results_forestplot(list(fs.est = fs), gbsg,
outcome.name = "rfstime", event.name = "status", treat.name = "hormon")
print(fp)
## End(Not run)
Print Method for K-Fold CV Results
Description
Print Method for K-Fold CV Results
Usage
## S3 method for class 'fs_kfold'
print(x, ...)
Arguments
x |
An fs_kfold object |
... |
Additional arguments (ignored) |
Value
x, invisibly. Called for its side effect of printing the
K-fold summary.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
oob <- forestsearch_Kfold(fs.est = fs)
print(oob)
## End(Not run)
Print method for fs_mr_oc objects
Description
Print method for fs_mr_oc objects
Usage
## S3 method for class 'fs_mr_oc'
print(x, ...)
Arguments
x |
An object of class |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print Method for fs_sg_plot Objects
Description
Print Method for fs_sg_plot Objects
Usage
## S3 method for class 'fs_sg_plot'
print(x, ...)
Arguments
x |
An fs_sg_plot object |
... |
Additional arguments (unused) |
Value
x, invisibly. Called for its side effect of printing the
subgroup definition, summary statistics and hazard ratio estimates.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
p <- plot_subgroup(fs)
print(p)
## End(Not run)
Print Method for Repeated K-Fold CV Results
Description
Print Method for Repeated K-Fold CV Results
Usage
## S3 method for class 'fs_tenfold'
print(x, ...)
Arguments
x |
An fs_tenfold object |
... |
Additional arguments (ignored) |
Value
x, invisibly. Called for its side effect of printing the
repeated K-fold summary.
Examples
## Not run:
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
kfold <- forestsearch_tenfold(fs.est = fs, sims = 5)
print(kfold)
## End(Not run)
Print Method for fs_weighted_km Objects
Description
Print Method for fs_weighted_km Objects
Usage
## S3 method for class 'fs_weighted_km'
print(x, ...)
Arguments
x |
An fs_weighted_km object from plot_sg_weighted_km() |
... |
Additional arguments (unused) |
Value
The input x, invisibly.
Print Method for gbsg_dgm Objects
Description
Print Method for gbsg_dgm Objects
Usage
## S3 method for class 'gbsg_dgm'
print(x, ...)
Arguments
x |
A gbsg_dgm object |
... |
Additional arguments (unused) |
Value
x, invisibly. Called for its side effect of printing the
data-generating mechanism summary.
Examples
## Not run:
dgm <- create_gbsg_dgm()
print(dgm)
## End(Not run)
Print method for glm_dgm objects
Description
Print method for glm_dgm objects
Usage
## S3 method for class 'glm_dgm'
print(x, ...)
Arguments
x |
A |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print method for glm_effect_profile objects
Description
Print method for glm_effect_profile objects
Usage
## S3 method for class 'glm_effect_profile'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments (unused). |
Value
Invisibly returns x.
Print method for guohe_a3 objects
Description
Prints the family size, the selected candidate, the naive and de-biased
effects, and the three interval objects returned by guohe_algorithm3(),
labelled by their status (Guo-He bound / extension / heuristic).
Usage
## S3 method for class 'guohe_a3'
print(x, ...)
Arguments
x |
An object of class |
... |
Ignored; present for method consistency. |
Value
x, invisibly.
Print method for guohe_ar objects
Description
Prints the cross-validation objective over the r grid, marks the selected
value, and, when present, prints the refitted full-data
guohe_algorithm3() result.
Usage
## S3 method for class 'guohe_ar'
print(x, ...)
Arguments
x |
An object of class |
... |
Ignored; present for method consistency. |
Value
x, invisibly.
Print method for hte_test objects
Description
Print method for hte_test objects
Usage
## S3 method for class 'hte_test'
print(x, ...)
Arguments
x |
An |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print method for leff_calibration objects
Description
Print method for leff_calibration objects
Usage
## S3 method for class 'leff_calibration'
print(x, ...)
Arguments
x |
A |
... |
Unused; present for S3 compatibility. |
Value
The input x, invisibly.
Print method for survreg_comparison objects
Description
Print method for survreg_comparison objects
Usage
## S3 method for class 'multi_survreg_comparison'
print(x, ...)
Arguments
x |
A survreg_comparison object |
... |
Additional arguments (not used) |
Value
Invisibly returns the input object
Examples
## Not run:
library(survival)
df <- survival::gbsg
res <- compare_multiple_survreg(df, outcome.name = "rfstime",
event.name = "status")
print(res)
## End(Not run)
Print CV ForestSearch Parameters
Description
Print CV ForestSearch Parameters
Usage
print_cv_params(cv_args)
Print detailed output for debugging
Description
Displays detailed information about the GRF analysis
Usage
print_grf_details(config, values, best_subgroup, sg_harm_id, tree_cuts = NULL)
Arguments
config |
List. GRF configuration |
values |
Data frame. Node metrics |
best_subgroup |
Data frame row. Selected subgroup (or NULL) |
sg_harm_id |
Character. Subgroup definition (or NULL) |
tree_cuts |
List. Cut information |
Value
Called for its side effect of printing to the console. The value
is returned invisibly and is incidental: tree_cuts$all when a subgroup,
its definition and tree_cuts are all supplied, otherwise NULL.
Examples
## Not run:
# print_grf_details() is called internally by grf.subg.harm.survival().
# See grf.subg.harm.survival() for the standard entry point.
## End(Not run)
Process forced cut expression for a variable
Description
Evaluates a cut expression (e.g., "age <= mean(age)") and returns the expression with the value.
Usage
process_conf_force_expr(expr, df)
Arguments
expr |
Character string of the cut expression. |
df |
Data frame. |
Value
Character string with evaluated value.
Examples
df <- data.frame(age = c(40, 55, 70, 35), size = c(20, 30, 25, 15))
process_conf_force_expr("age <= mean(age)", df)
process_conf_force_expr("size <= median(size)", df)
Process Continuous Variable for Subgroup Definition
Description
Process Continuous Variable for Subgroup Definition
Usage
process_continuous_subgroup(var_data, cut_spec, var_name, verbose)
Process Continuous Variables
Description
Process Continuous Variables
Usage
process_continuous_vars(
df_work,
data,
continuous_vars,
standardize,
marker = "z_"
)
Process Cutpoint Specification for Subgroup Definition
Description
Process Cutpoint Specification for Subgroup Definition
Usage
process_cutpoint(var_data, cut_spec, var_name = "", verbose = FALSE)
Process Factor Variable for Subgroup Definition
Description
Process Factor Variable for Subgroup Definition
Usage
process_factor_subgroup(var_data, cut_spec, var_name, verbose)
Process Factor Variables with Largest Value as Reference
Description
Process Factor Variables with Largest Value as Reference
Usage
process_factor_vars(df_work, data, factor_vars, marker = "z_")
75th Percentile (Quantile High)
Description
Returns the 75th percentile of a numeric vector.
Usage
qhigh(x)
Arguments
x |
A numeric vector. |
Value
Numeric value of the 75th percentile.
k-th J-Quantile
Description
Returns the (k/J)-th quantile of a numeric vector. Used by
cut_var_jq() to emit deferred cut expressions of the form
"X <= qj(X, k, J)", which are then resolved to literal
numerics by process_conf_force_expr().
Usage
qj(x, k, J)
Arguments
x |
A numeric vector. |
k |
Integer in |
J |
Integer >= 2. Total number of intervals implied by the
probability |
Value
Numeric value of the (k/J)-th quantile of x.
25th Percentile (Quantile Low)
Description
Returns the 25th percentile of a numeric vector.
Usage
qlow(x)
Arguments
x |
A numeric vector. |
Value
Numeric value of the 25th percentile.
Quick Plot KM Bands from ForestSearch
Description
Convenience wrapper with sensible defaults for quick visualization.
Usage
quick_km_band_plot(df, fs.est, outcome.name, event.name, treat.name, ...)
Arguments
df |
Data frame with analysis data. |
fs.est |
ForestSearch result object. |
outcome.name |
Character. Time-to-event column name. |
event.name |
Character. Event indicator column name. |
treat.name |
Character. Treatment column name. |
... |
Additional arguments passed to |
Value
Invisibly returns the plot result.
Examples
## Not run:
# Quick plot with minimal configuration
quick_km_band_plot(df_analysis, fs_result, "os_months", "os_event", "treat")
# With reference subgroups
ref_sgs <- list(
age_young = list(subset_expr = "age < 65", color = "brown"),
age_old = list(subset_expr = "age >= 65", color = "grey")
)
quick_km_band_plot(df_analysis, fs_result, "os_months", "os_event", "treat",
ref_subgroups = ref_sgs, tau_add = 60)
## End(Not run)
Remove Near-Duplicate Subgroups
Description
Removes subgroups with nearly identical statistics (HR, n, E, etc.) to reduce redundancy in candidate list.
Usage
remove_near_duplicate_subgroups(
hr_subgroups,
tolerance = 0.001,
details = FALSE
)
Arguments
hr_subgroups |
Data.table of subgroup results. |
tolerance |
Numeric. Tolerance for numeric comparison (default 0.001). |
details |
Logical. Print details about removed duplicates. |
Value
Data.table with near-duplicate rows removed.
Remove Redundant Subgroups
Description
Removes redundant subgroups by checking for exact ties in key columns.
Usage
remove_redundant_subgroups(found.hrs)
Arguments
found.hrs |
Data.table of found subgroups. |
Value
Data.table of non-redundant subgroups.
Render ForestSearch Forest Plot
Description
Renders a forest plot from plot_subgroup_results_forestplot().
Usage
render_forestplot(x, newpage = TRUE)
Arguments
x |
An fs_forestplot object from |
newpage |
Logical. Call grid.newpage() before drawing. Default: TRUE. |
Details
To control plot sizing, create a custom theme using create_forest_theme()
and pass it to plot_subgroup_results_forestplot():
my_theme <- create_forest_theme(base_size = 14, row_padding = c(6, 4))
result <- plot_subgroup_results_forestplot(..., theme = my_theme)
render_forestplot(result)
Value
Invisibly returns the grob object.
Examples
## Not run:
# For larger plot, use theme parameter in plot_subgroup_results_forestplot:
large_theme <- create_forest_theme(
base_size = 14,
row_padding = c(6, 4),
cv_fontsize = 11
)
result <- plot_subgroup_results_forestplot(
fs_results = list(fs.est = fs, fs_bc = fs_bc),
df_analysis = df.analysis,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
theme = large_theme
)
render_forestplot(result)
## End(Not run)
Render Reference Simulation Table as gt
Description
Converts a data frame of pre-computed reference simulation results (e.g.,
digitized from a published LaTeX table) into a styled gt table. This
is useful for displaying published benchmark results alongside new
simulation output within vignettes or reports.
Usage
render_reference_table(
ref_df,
title = "Reference Simulation Results",
subtitle = NULL,
bold_threshold = 0.05,
subgroup_notation = c("harm", "benefit")
)
Arguments
ref_df |
|
title |
Character. Table title. |
subtitle |
Character. Table subtitle. Default: |
bold_threshold |
Numeric. Values in |
subgroup_notation |
Character. |
Value
A gt table object.
Examples
## Not run:
ref <- data.frame(
Scenario = "M1 Null: N=700",
Metric = "any(H)",
FS = 0.02,
FSlg = 0.03,
GRF = 0.25
)
render_reference_table(ref, title = "Reference Results")
## End(Not run)
Reset a Parallel Worker Pool
Description
Tears down the current future plan, forces garbage collection, optionally pauses, and spins up a fresh worker pool. Used to release worker-side memory between simulation phases (e.g. between alternative and null Monte Carlo runs) so that accumulated R-session state in long-running workers does not bloat memory or cause sporadic failures.
Usage
reset_workers(
workers = NULL,
strategy = "multisession",
sleep = 1,
gc = TRUE,
quiet = FALSE
)
Arguments
workers |
Integer. Number of workers for the new pool.
Default |
strategy |
Character. The future strategy to
reinstate. Default |
sleep |
Numeric. Seconds to pause between teardown and
setup. Default |
gc |
Logical. Whether to force garbage collection between
teardown and setup. Default |
quiet |
Logical. Suppress informational |
Details
The worker reset sequence is:
-
future::plan("sequential")– dismantles the existing worker pool; R processes terminate and their memory is returned to the OS. -
gc(verbose = FALSE)– reclaims main-process memory holding references to now-defunct futures. -
Sys.sleep(sleep)– brief pause to let the OS reclaim sockets / pipes before the new pool is requested. A small non-zero sleep (the default1) materially reducesmultisessionsetup failures observed on busy Linux servers; set to0to skip. -
future::plan(strategy, workers = workers)– starts a fresh pool.
This function is a no-op if future is not installed or if
the requested strategy cannot be initialised; a warning is emitted
and the previous plan is left in place.
Value
Invisibly, a list with the previous and current plan
class names: list(previous = <chr>, current = <chr>).
See Also
collect_results for error-tolerant
collection of foreach output.
Examples
if (requireNamespace("future", quietly = TRUE)) {
future::plan("multisession", workers = 2L)
# ... run alternative simulations ...
reset_workers(workers = 2L)
# ... run null simulations ...
future::plan("sequential")
}
Resolve parallel processing arguments for bootstrap
Description
If parallel_args not provided, falls back to forestsearch call's parallel configuration. Always reports configuration to user.
Usage
resolve_bootstrap_parallel_args(parallel_args, forestsearch_call_args)
Arguments
parallel_args |
List or empty list |
forestsearch_call_args |
List from original forestsearch call |
Value
List with plan, workers, show_message
Resolve Parallel Arguments for Cross-Validation
Description
Helper function to resolve and validate parallel processing arguments,
similar to bootstrap's resolve_bootstrap_parallel_args.
Usage
resolve_cv_parallel_args(parallel_args, fs_args, details = FALSE)
Arguments
parallel_args |
List. User-provided parallel arguments. |
fs_args |
List. Original ForestSearch call arguments. |
details |
Logical. Print configuration messages. |
Value
List with resolved plan, workers, show_message.
RMST calculation for subgroup
Description
Calculates restricted mean survival time (RMST) for a subgroup.
Usage
rmst_calculation(
df,
tte.name = "tte",
event.name = "event",
treat.name = "treat"
)
Arguments
df |
Data frame. |
tte.name |
Character. Name of time-to-event variable. |
event.name |
Character. Name of event indicator variable. |
treat.name |
Character. Name of treatment variable. |
Value
List with tau, RMST, RMST for treatment, RMST for control.
Examples
## Not run:
library(survival)
df <- data.frame(
tte = gbsg$rfstime / 30.4375,
event = gbsg$status,
treat = gbsg$hormon
)
rmst_calculation(df, tte.name = "tte", event.name = "event",
treat.name = "treat")
## End(Not run)
Run Null Calibration Simulation
Description
Runs ForestSearch under the null hypothesis at multiple sample
sizes to calibrate L_{\text{eff}}, then returns a
calibration object for predicting corrected FPR.
Usage
run_null_calibration(
sim_null_fn,
fs_args,
d_eff_fn,
N_values = c(200, 500, 1000),
c1 = 1.25,
c2 = 1,
n_sims = 100L,
prop_harm = 0.3,
n_min = 60L,
seed_base = 42L,
verbose = TRUE
)
Arguments
sim_null_fn |
Function. Generates one H0 dataset.
Signature: |
fs_args |
Named list. Arguments passed to |
d_eff_fn |
Function. Computes |
N_values |
Integer vector. Sample sizes to simulate.
At least 2 required. Default: |
c1 |
Numeric. Screening threshold for calibration. Default: 1.25. |
c2 |
Numeric. Consistency threshold for calibration. Default: 1.0. |
n_sims |
Integer. Simulations per sample size. Default: 100. |
prop_harm |
Numeric. Expected proportion in harm subgroup (for computing d_eff). Default: 0.30. |
n_min |
Integer. Minimum subgroup size. Default: 60. |
seed_base |
Integer. Base seed. Default: 42. |
verbose |
Logical. Print progress. Default: TRUE. |
Details
For each N in N_values, the function:
Generates
n_simsdatasets viasim_null_fn()Runs
forestsearch()on eachComputes the simulated FPR = proportion with any H found
Computes
P_1via the GLM approximationCalls
calibrate_L_eff()to fit the power-law model
ForestSearch is run with
parallel_args = list(plan = "sequential") to avoid
nested parallelism. The outer simulation loop is sequential.
For faster execution, wrap in foreach %dofuture%
externally.
Value
An "leff_calibration" object.
Examples
# Binary DGM under H0
sim_null <- function(n, seed) {
set.seed(seed)
data.frame(
id = seq_len(n),
treat = rbinom(n, 1, 0.5),
bm1 = as.factor(rbinom(n, 1, 0.70)),
bm2 = as.factor(rbinom(n, 1, 0.50)),
age = round(rnorm(n, 55, 10)),
ecog = as.factor(sample(0:1, n, TRUE, c(0.6, 0.4))),
progressed = rbinom(n, 1, 0.30)
)
}
fs_args <- list(
confounders.name = c("bm1", "bm2", "age", "ecog"),
outcome.name = "progressed",
event.name = "progressed",
treat.name = "treat",
id.name = "id",
outcome_type = "binary",
effect_measure = "OR",
adverse_outcome = TRUE,
pconsistency.threshold = 0.90,
fs.splits = 200,
n.min = 60,
maxk = 2,
use_lasso = TRUE,
use_grf = TRUE,
use_twostage = TRUE,
is.RCT = TRUE,
details = FALSE,
plot.sg = FALSE,
parallel_args = list(plan = "sequential",
workers = 1,
show_message = FALSE)
)
cal <- run_null_calibration(
sim_null_fn = sim_null,
fs_args = fs_args,
d_eff_fn = function(n) d_eff_binary(n, 0.30),
N_values = c(200, 500, 1000),
n_sims = 50,
verbose = TRUE
)
print(cal)
Run One Simulation Replicate
Description
General replacement for the legacy run_simulation_analysis() that
was coupled to simulate_from_gbsg_dgm() and GBSG-specific column
names. This version calls simulate_from_dgm and accepts
explicit column-name parameters, making it applicable to any DGM built
with generate_aft_dgm_flex.
Usage
run_simulation_analysis(
sim_id,
dgm,
n_sample,
analysis_time = Inf,
cens_adjust = 0,
max_follow = NULL,
muC_adj = NULL,
confounders_base = c("v1", "v2", "v3", "v4", "v5", "v6", "v7"),
n_add_noise = 0L,
outcome_name = "y_sim",
event_name = "event_sim",
treat_name = "treat_sim",
harm_col = "flag_harm",
run_fs = TRUE,
run_fs_grf = TRUE,
run_grf = TRUE,
methods = NULL,
fs_params = list(),
grf_params = list(),
cox_formula = NULL,
cox_formula_adj = NULL,
n_sims_total = NULL,
seed_base = 8316951L,
verbose = FALSE,
verbose_n = NULL,
debug = FALSE,
keep = character(0),
keep_first_n = Inf
)
Arguments
sim_id |
Integer. Simulation replicate index (used as seed offset). |
dgm |
An |
n_sample |
Integer. Per-replicate sample size. |
analysis_time |
Numeric. Calendar time of analysis on the DGM time
scale. Use |
cens_adjust |
Numeric. Log-scale shift to censoring times passed to
|
max_follow |
Deprecated. Use |
muC_adj |
Deprecated. Use |
confounders_base |
Character vector of base confounder names. |
n_add_noise |
Integer. Number of independent N(0,1) noise variables
to append. Default |
outcome_name |
Name of the observed time column in simulated data.
Default |
event_name |
Name of the event indicator column. Default
|
treat_name |
Name of the treatment column. Default |
harm_col |
Name of the true-subgroup indicator column. Default
|
run_fs |
Logical. Run ForestSearch (LASSO). Default |
run_fs_grf |
Logical. Run ForestSearch (LASSO + GRF). Default
|
run_grf |
Logical. Run standalone GRF. Default |
methods |
Optional specification of analysis arms that generalizes the
Each arm's overrides are merged on top of |
fs_params |
Named list of ForestSearch parameter overrides. Any
element of |
grf_params |
Named list of GRF parameter overrides. |
cox_formula |
Optional Cox formula for unadjusted ITT. |
cox_formula_adj |
Optional adjusted Cox formula. |
n_sims_total |
Integer. Total simulations (for progress messages). |
seed_base |
Integer. Base seed; replicate seed = |
verbose |
Logical. Print progress messages. Default |
verbose_n |
Integer. If set, only print for |
debug |
Logical. Print detailed debug output. Default |
keep |
Character vector. Optional names of heavy diagnostic
objects to attach as a list in
Default |
keep_first_n |
Integer. When |
Value
A data.table with one row per analysis method.
Always-present scalar columns cover: true-subgroup size/effect
from the DGM, per-method detection flag (any.H), estimated
subgroup size (size.H, size.Hc), effect estimates
(hr.H.hat, hr.Hc.hat, hr.itt, and their true
counterparts), classification metrics against the DGM-stored
subgroup (sens, spec, ppv, npv), and
pairwise concordance between methods (agree.*.*,
kappa.*.*). As of v0.2.0 the following ForestSearch
diagnostic columns are also always populated (NA for GRF
rows):
sg.defCharacter. Cut-expression string of the identified subgroup, formed as
paste(fs_full$sg.harm, collapse = " & "). Empty string when no subgroup was identified. Read across all replicates to diagnose which cuts the method is selecting (e.g.,table(results$sg.def)).sg.n_factorsInteger. Number of factors in the identified subgroup definition.
n_candidates_evaluatedInteger. Number of candidate subgroups actually evaluated for consistency (before or at the early-stop candidate).
n_candidates_totalInteger. Total number of candidate subgroups available.
n_passedInteger. Number of candidates meeting the consistency threshold.
consistency_algorithmCharacter.
"fixed"or"twostage".early_stop_triggeredLogical. Whether evaluation stopped early due to
stop_threshold.fs_minutesNumeric. Wall time for the FS run in minutes.
The analogous GRF row carries sg.def (from
grf_result$sg.harm.id) and grf_selected_depth.
When keep is non-empty and sim_id <= keep_first_n,
the heavy objects are attached as
attr(result, "diagnostics") — a named list with one
element per keeper (fs_full, grf_full, etc.).
GLM Parameters
GLM-specific parameters (outcome_type, effect_measure,
offset.name) must be passed inside fs_params, not as
top-level arguments. They route only to the estimation step
(.extract_fs_estimates_gen, .extract_grf_estimates_gen),
not to forestsearch() itself, which uses Cox PH for subgroup
identification in v0.1.x. Passing these as top-level arguments will
result in them being silently ignored.
Parallel Processing
run_simulation_analysis() is designed to be called inside a
foreach() %dofuture% loop (one replicate per worker). In
that idiom, outer parallelism is provided by %dofuture%
(one replicate per worker) and the inner forestsearch()
call should run sequentially within each worker. Running both layers
multisession produces nested parallelism: each outer worker tries to
spawn its own pool of inner workers, which parallelly rejects
with a 300\
To prevent this, default_sim_params() sets
parallel_args = list(plan = "sequential") as the default for
the inner forestsearch() call. Users who call
run_simulation_analysis() once at the top level (e.g., for
interactive debugging) and want the inner pipeline to run in parallel
can opt back in by passing
fs_params = list(parallel_args = list(plan = "multisession", workers = N)).
See Also
simulate_from_dgm,
generate_aft_dgm_flex, setup_gbsg_dgm
Examples
## Not run:
dgm <- setup_gbsg_dgm(model = "null", verbose = FALSE)
# Inner forestsearch() runs sequentially by default; safe for both
# interactive use and embedding inside foreach() %dofuture% loops.
result <- run_simulation_analysis(sim_id = 1, dgm = dgm, n_sample = 500)
# Opt back into inner multisession for a one-off top-level call:
result <- run_simulation_analysis(
sim_id = 1,
dgm = dgm,
n_sample = 500,
fs_params = list(parallel_args = list(plan = "multisession",
workers = 4))
)
## End(Not run)
Run Single Consistency Split
Description
Performs one random 50/50 split and evaluates whether both halves meet the HR consistency threshold.
Usage
run_single_consistency_split(
df.x,
N.x,
hr.consistency,
cox_init = 0,
estimator_fn = NULL,
consistency_threshold = NULL,
adjust_covariates = NULL
)
Arguments
df.x |
data.table. Subgroup data with columns Y, Event, Treat. |
N.x |
Integer. Number of observations in subgroup. |
hr.consistency |
Numeric. |
cox_init |
Numeric. Initial value for Cox model (log HR). |
estimator_fn |
Closure or |
consistency_threshold |
Numeric or |
adjust_covariates |
Character vector or |
Value
Numeric. 1 if both splits meet threshold, 0 if not, NA if error.
Examples
library(data.table)
set.seed(1)
df <- data.table(
Y = rexp(100),
Event = rbinom(100, 1, 0.55),
Treat = rep(0:1, 50)
)
run_single_consistency_split(df, N.x = 100, hr.consistency = 1.0)
Run the extreme-subgroups simulation study
Description
Draws n_sims synthetic trials from a calibrated DGM via
simulate_from_dgm(), fits fit in every pre-defined subgroup of
every trial, and returns the raw result matrices plus full
provenance. With the default fit, benchmarks, seeds, and skip
rule, the matrices are bit-identical to those produced by the
extreme_subgroups vignettes for the matching baseline, and the
returned object is byte-compatible with their RDS payload:
saveRDS(sims, path) yields a file the design-comparison memo loads
unchanged.
Usage
run_subgroup_sims(
dgm,
subgroups,
n_sims,
fit = subgroup_cox(survival::Surv(y_sim, event_sim) ~ treat_sim),
baseline = c("resample", "fixed"),
n = NULL,
analysis_time,
max_entry,
cens_adjust,
cutpoints = list(),
benchmarks = benchmark_spec(),
min_n = 5L,
workers = NULL,
seed_base = 0L,
rand_seed_offset = 1000000L,
hr_true = NULL,
k_treat = NULL,
future_packages = c("forestsearch", "survival"),
validate = TRUE,
verbose = FALSE
)
Arguments
dgm |
A DGM from |
subgroups |
List of subgroup definitions, each a list with
character scalars |
n_sims |
Number of simulated trials. |
fit |
Per-subgroup analysis function |
baseline |
|
n |
Patients per trial. Required for |
analysis_time, max_entry, cens_adjust |
Passed to
|
cutpoints |
Named list of scalars exposed as columns of each
simulated trial so |
benchmarks |
A |
min_n |
Subgroups with fewer rows are skipped for that trial
(row stays |
workers |
Parallel workers. |
seed_base, rand_seed_offset |
Seed scheme; defaults ( |
hr_true, k_treat |
Optional provenance values stored verbatim in
the result (the vignettes store |
future_packages |
Packages loaded on workers, default
|
validate |
Run |
verbose |
Print progress lines. |
Details
Per-iteration seeding is explicit (trial ss uses
seed_base + ss for the DGM draw and
seed_base + ss + rand_seed_offset for benchmark membership), so
results are independent of the parallel schedule: any workers
setting, including sequential, gives identical matrices.
GLM dispatch: when dgm inherits "glm_dgm" (generate_glm_dgm()),
trials are drawn by simulate_from_glm_dgm() (n is required; the
survival-only analysis_time / max_entry / cens_adjust must not
be supplied; baseline = "fixed" uses the stored df_source panel), the default
fit becomes subgroup_glm() constructed from the DGM's own
outcome_type and effect_measure (binary DGMs also pass their
adverse_outcome; a continuous DGM resolves to the same fitter as
before), and the result additionally carries
the fitter's effect scale metadata with design = "glm". The
sim_hrs / sim_ubs matrices then hold that fitter's
(estimate, upper bound) pairs – the field names are retained for
structural compatibility, exactly as hr.threshold serves
generic-threshold duty for GLM outcomes in forestsearch().
Value
An object of class "subgroup_sims": a named list with
design, n_sims, matrices sim_hrs / sim_ubs / sim_ns
(n_sims x subgroups, columns named by subgroup name),
subgroups, sim_config, cens_adjust, k_treat, hr_true,
created, r_version, and forestsearch_version. For a
"glm_dgm": design = "glm", cens_adjust = NULL, and an
additional effect field holding the fitter's scale metadata (see
subgroup_glm()); survival results are constructed exactly as
before, with no new fields.
See Also
summary.subgroup_sims(), benchmark_spec(),
subgroup_cox(), subgroup_glm(), validate_subgroups()
Examples
## Not run:
sims <- run_subgroup_sims(
dgm = dgm_uniform, subgroups = subgroups, n_sims = 1000,
fit = subgroup_cox(survival::Surv(y_sim, event_sim) ~ treat_sim +
survival::strata(grade)),
baseline = "fixed",
analysis_time = 84, max_entry = 24,
cens_adjust = cal_uniform$cens_adjust,
cutpoints = list(age_med = median(gbsg$age)),
hr_true = 0.70, k_treat = k_treat_uniform
)
saveRDS(sims, "results/extreme_sims_fixed_1000_payload.rds")
## End(Not run)
Evaluate an expression string in a data-frame scope
Description
Parses and evaluates expr in a restricted environment
containing only the columns of df (parent: baseenv()).
This isolates evaluation from the global environment, reducing
scope for unintended side effects.
Usage
safe_eval_expr(df, expr)
Arguments
df |
Data frame providing column names as variables. |
expr |
Character. Expression to evaluate
(e.g., |
Value
Result of evaluating expr, or NULL on failure.
Note
eval(parse()) is used intentionally here.
evaluate_comparison handles only single comparisons
(e.g., "er <= 0"); this function is needed for the compound
logical expressions produced by the ForestSearch subgroup enumeration
algorithm (e.g., "er <= 0 & nodes > 3"). Evaluation is
sandboxed: the environment contains only the columns of df
with baseenv() as parent, so neither the global environment
nor any package namespace is in scope. No user-supplied strings are
evaluated; only internally-constructed subgroup definition strings
reach this function.
See Also
evaluate_comparison for the single-comparison
operator-dispatch alternative that avoids eval(parse()).
Subset a data frame using an expression string
Description
Thin wrapper around safe_eval_expr that uses the
logical result to subset rows.
Usage
safe_subset(df, expr)
Arguments
df |
Data frame. |
expr |
Character. Subset expression
(e.g., |
Value
Subset of df, or NULL on failure.
GLM Sandwich Variance
Description
GLM Sandwich Variance
Usage
sandwich_var_glm(P, Y, mu, family)
OLS Sandwich Variance (HC0)
Description
OLS Sandwich Variance (HC0)
Usage
sandwich_var_ols(P, Y, beta_hat)
Save ForestSearch Forest Plot to File
Description
Saves a forest plot to a file (PDF, PNG, etc.) with explicit dimensions.
Usage
save_forestplot(x, filename, width = 12, height = 10, dpi = 300, bg = "white")
Arguments
x |
An fs_forestplot object. |
filename |
Character. Output filename. Extension determines format. |
width |
Numeric. Plot width in inches. Default: 12. |
height |
Numeric. Plot height in inches. Default: 10. |
dpi |
Numeric. Resolution for raster formats. Default: 300. |
bg |
Character. Background color. Default: "white". |
Value
Invisibly returns the filename.
Examples
## Not run:
# Create plot with custom theme
large_theme <- create_forest_theme(base_size = 14, row_padding = c(6, 4))
result <- plot_subgroup_results_forestplot(
fs_results = list(fs.est = fs, fs_bc = fs_bc),
df_analysis = df.analysis,
outcome.name = "time",
event.name = "status",
treat.name = "treatment",
theme = large_theme
)
# Save to file
save_forestplot(result, "forest_plot.pdf", width = 14, height = 12)
## End(Not run)
Save L_eff Calibration Results
Description
Writes a standardized .rds bundle containing the fitted
L_eff parameters and all supporting data. Called at the end of
calibration documents (e.g., calibration_binary_leff.qmd).
Usage
save_leff_calibration(
C,
alpha,
n_min,
outcome_type,
cal_data = NULL,
dgm_description = "",
n_sims_per_N = NA_integer_,
output_dir = NULL,
extra = list()
)
Arguments
C |
Numeric. Fitted scale parameter. |
alpha |
Numeric. Fitted power-law exponent. |
n_min |
Integer. Reference minimum subgroup size. |
outcome_type |
Character. One of "binary", "survival", "count", "continuous". |
cal_data |
Data frame with columns N, sim_fpr, P1, L_eff (the raw calibration points). |
dgm_description |
Character. Free-text description of the DGM used for calibration. |
n_sims_per_N |
Integer. Simulations per sample size. |
output_dir |
Character. Directory for .rds files. Required: an
error is raised when it is |
extra |
List. Any additional items to store. |
Value
Invisible path to the saved file.
Examples
## Not run:
save_leff_calibration(
C = 0.220, alpha = 1.298, n_min = 60,
outcome_type = "binary",
cal_data = cal_df,
dgm_description = "Binary DGM, 4 confounders (bm1, bm2, age, ecog)",
n_sims_per_N = 5000,
output_dir = tempdir()
)
## End(Not run)
Select best subgroup based on criterion
Description
Identifies the optimal subgroup according to the specified criterion
Usage
select_best_subgroup(values, sg.criterion, dmin.grf, n.max)
Arguments
values |
Data frame. Node metrics from policy trees |
sg.criterion |
Character. "mDiff" for maximum difference, "Nsg" for largest size |
dmin.grf |
Numeric. Minimum difference threshold |
n.max |
Integer. Maximum allowed subgroup size (total sample size) |
Value
Data frame row with best subgroup or NULL if none found
Examples
vals <- data.frame(diff = c(8.5, 6.2, 3.1), Nsg = c(120, 95, 80))
select_best_subgroup(values = vals, sg.criterion = "mDiff",
dmin.grf = 6, n.max = 500)
Generate Cross-Validation Sensitivity Text
Description
Creates formatted text summarizing cross-validation agreement metrics.
Usage
sens_text(fs_kfold, est.scale = "hr")
Arguments
fs_kfold |
K-fold cross-validation results from forestsearch_Kfold. |
est.scale |
Character. "hr" or "1/hr". |
Value
Character string with formatted CV metrics.
Sensitivity Analysis of Hazard Ratios to k_inter
Description
Analyzes how the interaction parameter k_inter affects hazard ratios in different populations (overall, harm subgroup, no-harm subgroup).
Usage
sensitivity_analysis_k_inter(
k_inter_range = c(-5, 5),
n_points = 21,
plot = TRUE,
...
)
Arguments
k_inter_range |
Numeric vector of length 2 specifying the range of k_inter values to analyze. Default is c(-5, 5). |
n_points |
Integer number of points to evaluate within the range. Default is 21. |
plot |
Logical indicating whether to create visualization plots. Default is TRUE. |
... |
Additional arguments passed to |
Details
This function evaluates the hazard ratios at evenly spaced points across the k_inter range. If plot = TRUE, it creates a 4-panel visualization showing:
Harm subgroup HR vs k_inter
All HRs (overall, harm, no-harm) vs k_inter
Ratio of HRs (harm/no-harm) showing effect modification
Table of key values
Value
A data.frame of class "k_inter_sensitivity" with columns:
- k_inter
Numeric k_inter value
- hr_harm
Numeric hazard ratio in harm subgroup
- hr_no_harm
Numeric hazard ratio in no-harm subgroup
- hr_overall
Numeric overall hazard ratio
- subgroup_size
Integer size of harm subgroup
Examples
## Not run:
# Analyze sensitivity to k_inter
sensitivity_results <- sensitivity_analysis_k_inter(
k_inter_range = c(-2, 2),
n_points = 11,
data = survival::gbsg,
continuous_vars = c("age", "er", "pgr"),
factor_vars = c("meno", "grade"),
outcome_var = "rfstime",
event_var = "status",
treatment_var = "hormon",
subgroup_vars = c("er", "meno"),
subgroup_cuts = list(er = 20, meno = 0),
model = "alt",
plot = TRUE
)
# Results show relationship between k_inter and HRs
print(sensitivity_results)
## End(Not run)
Set Up a GBSG-Based AFT Data Generating Mechanism
Description
Creates a GBSG-based data generating mechanism that is fully compatible with
simulate_from_dgm. This is the replacement for
create_gbsg_dgm(): it accepts exactly the same arguments and produces
the same numeric output, but returns an object of class
"aft_dgm_flex" instead of "gbsg_dgm".
Usage
setup_gbsg_dgm(
model = c("alt", "null"),
k_treat = 1,
k_inter = 1,
k_z3 = 1,
z1_quantile = 0.25,
n_super = 5000L,
cens_type = c("weibull", "uniform"),
use_rand_params = FALSE,
seed = 8316951L,
verbose = FALSE,
k_random_noise = 0L,
noise_seed = 20260807L
)
Arguments
model |
Character. Either "alt" for alternative hypothesis with heterogeneous treatment effects, or "null" for uniform treatment effect. Default: "alt" |
k_treat |
Numeric. Treatment effect multiplier applied to the treatment coefficient from the fitted AFT model. Values > 1 strengthen the treatment effect. Default: 1 |
k_inter |
Numeric. Interaction effect multiplier for the treatment-subgroup interaction (z1 * z3). Only used when model = "alt". Higher values create more heterogeneity between HR(H) and HR(Hc). Default: 1 |
k_z3 |
Numeric. Effect multiplier for the z3 (menopausal status) coefficient. Default: 1 |
z1_quantile |
Numeric. Quantile threshold for z1 (estrogen receptor). Observations with ER <= quantile are coded as z1 = 1. Default: 0.25 |
n_super |
Integer. Size of super-population for empirical HR estimation. Default: 5000 |
cens_type |
Character. Censoring distribution type: "weibull" or "uniform". Default: "weibull" |
use_rand_params |
Logical. If TRUE, modifies confounder coefficients using estimates from randomized subset (meno == 0). Default: FALSE |
seed |
Integer. Random seed for super-population generation. Default: 8316951 |
verbose |
Logical. Print diagnostic information. Default: FALSE |
k_random_noise |
Integer. Number of standard-normal noise columns
( |
noise_seed |
Integer. Seed for the population-level noise draw, used
only when |
Details
Internally the function calls create_gbsg_dgm() and then:
Adds a
df_superfield with column names aligned tosimulate_from_dgm()conventions (lin_pred_1,lin_pred_0,lin_pred_cens_1,lin_pred_cens_0,flag_harm).Adds a
model_params$taufield (=model_params$sigma) and amodel_params$censoringsub-list.Sets class to
c("aft_dgm_flex", "gbsg_dgm", "list").
The original df_super_rand field is kept so that
compute_dgm_cde() and print.gbsg_dgm continue to work.
Noise columns requested via k_random_noise are inert
N(0,1) population attributes intended as candidate confounders; the
builder never puts them in the outcome model. They are drawn once onto
the super-population (both df_super_rand and df_super),
so simulated trials and evaluation frames inherit them by row resampling.
Value
An object of class c("aft_dgm_flex", "gbsg_dgm", "list")
with all fields from create_gbsg_dgm() plus:
df_superSuper-population data frame with
simulate_from_dgm()-compatible column names.model_params$tauCopy of
model_params$sigma.model_params$censoringSub-list with
type,mu,taufor the censoring model.noise_namesCharacter vector of noise column names (
character(0)whenk_random_noise = 0).noise_scheme"population"when noise was drawn, otherwise"none".noise_seedThe seed used for the noise draw, or
NA_integer_when none was drawn.
See Also
create_gbsg_dgm, simulate_from_dgm,
compute_dgm_cde
Examples
## Not run:
dgm <- setup_gbsg_dgm(model = "alt", k_inter = 2, verbose = FALSE)
dgm <- compute_dgm_cde(dgm)
print(dgm)
sim <- simulate_from_dgm(dgm, n = 400, seed = 1)
## End(Not run)
Set up parallel processing for subgroup consistency
Description
Sets up parallel processing using the specified approach and number of workers.
Usage
setup_parallel_SGcons(
parallel_args = list(plan = "multisession", workers = 4, show_message = TRUE)
)
Arguments
parallel_args |
List with |
Value
None. Sets up parallel backend as side effect.
Examples
setup_parallel_SGcons(list(plan = "sequential", workers = 1,
show_message = FALSE))
future::plan(future::sequential) # reset
Output Subgroup Consistency Results
Description
Returns the top subgroup(s) and recommended treatment flags.
Usage
sg_consistency_out(
df,
result_new,
sg_focus,
index.Z,
names.Z,
details = FALSE,
plot.sg = FALSE,
by.risk = 12,
confs_labels,
is_glm = FALSE,
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
effect_log_scale = FALSE
)
Arguments
df |
Data.frame. Original analysis data. |
result_new |
Data.table. Sorted subgroup results. |
sg_focus |
Character. Sorting focus criterion. |
index.Z |
Matrix. Subgroup factor indicators. |
names.Z |
Character vector. Factor column names. |
details |
Logical. Print details. |
plot.sg |
Logical. Plot subgroup curves. |
by.risk |
Numeric. Risk interval for plotting. |
confs_labels |
Character vector. Human-readable labels. |
is_glm |
Logical. If |
selection_rule |
Character. Rule defining the candidate
inclusion set for |
effect_neighborhood |
Numeric in |
effect_log_scale |
Logical. If |
Value
List with elements:
- result
Sorted candidate table (top row = selected subgroup).
- pareto_frontier
Data.table of non-dominated candidates on (effect, N), both maximized. See
compute_pareto_frontier. May beNULLif computation failed.- sg.harm
Factor-level cut names defining the selected subgroup.
- sg.harm_label
Human-readable subgroup labels.
- df_flag
Per-subject treatment-recommendation flags.
- sg.harm.id
Per-subject subgroup-membership indicator.
Examples
## Not run:
# sg_consistency_out is called internally by forestsearch().
# See forestsearch() for the standard entry point.
fs <- forestsearch(gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
fs$grp.consistency$out_sg
## End(Not run)
Enhanced Subgroup Summary Tables (gt output)
Description
Returns formatted summary tables for subgroups using the gt package, with search metadata and customizable decimal precision. Produces two tables: a treatment effect estimates table and an identified subgroups table, each with fully customizable titles and subtitles.
Usage
sg_tables(
fs,
which_df = "est",
est_title = "Treatment Effect Estimates",
est_caption = "Training data estimates",
sg_title = "Identified Subgroups",
sg_subtitle = NULL,
potentialOutcome.name = NULL,
hr_1a = NA,
hr_0a = NA,
ndecimals = 3,
include_search_info = TRUE,
subgroup_notation = NULL,
font_size = 12
)
Arguments
fs |
ForestSearch results object, or a |
which_df |
Character. Which data frame to use ("est" or "testing"). |
est_title |
Character or NULL. Main title for the estimates table
(default: "Treatment Effect Estimates"). Rendered as bold markdown.
Set to NULL to suppress the title and display only |
est_caption |
Character. Subtitle for the estimates table (default: "Training data estimates"). |
sg_title |
Character or NULL. Main title for the identified subgroups
table (default: "Identified Subgroups"). Rendered as bold markdown.
Set to NULL to suppress the title and display only |
sg_subtitle |
Character or NULL. Subtitle for the identified subgroups
table. When NULL (default), an informative subtitle is auto-generated
from |
potentialOutcome.name |
Character. Name of potential outcome variable (optional). |
hr_1a |
Character. Adjusted HR for subgroup 1 (optional). |
hr_0a |
Character. Adjusted HR for subgroup 0 (optional). |
ndecimals |
Integer. Number of decimals for formatted numbers
(default: 3). Controls precision in both |
include_search_info |
Logical. Include search metadata table (default: TRUE). |
subgroup_notation |
Character or |
font_size |
Numeric. Font size in pixels for table text (default: 12). |
Value
List with gt tables for estimates, subgroups, and optionally search info.
Examples
## Not run:
library(survival)
fs <- forestsearch(
gbsg,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime",
treat.name = "hormon",
event.name = "status",
fs.splits = 50,
use_lasso = FALSE,
use_grf = FALSE
)
tabs <- sg_tables(fs)
tabs$tab_estimates
## End(Not run)
Adaptive Figure Height for a Subgroup-Distribution Plot
Description
Computes a reasonable knitr fig.height (in inches) for a
plot_sg_distribution() panel based on the number of
bars, keeping the visual quality consistent across scenarios with
very different numbers of distinct subgroup labels.
Usage
sgdist_fig_height(
n_bars,
floor_in = 4,
ceil_in = 10,
per_bar = 0.35,
base_in = 3
)
Arguments
n_bars |
Integer. Number of bars in the plot, typically
retrieved via |
floor_in |
Numeric. Minimum height in inches.
Default: |
ceil_in |
Numeric. Maximum height in inches.
Default: |
per_bar |
Numeric. Height contribution per bar in inches.
Default: |
base_in |
Numeric. Base height added to |
Value
Numeric scalar, a height in inches in [floor_in,
ceil_in]. Suitable for the fig.height chunk option.
See Also
Examples
## Not run:
p <- plot_sg_distribution(results, placeholder_on_empty = TRUE)
h <- sgdist_fig_height(attr(p, "n_bars"))
# Use `h` as the fig.height chunk option.
## End(Not run)
Region-consistency metrics on a simulated trial (ITT and optionally a subgroup)
Description
Region-consistency metrics on a simulated trial (ITT and optionally a subgroup)
Usage
sim_region_metrics(sim, region_col, subgroup_index = NULL, pi_preserve = 0.5)
Arguments
sim |
Data frame of one simulated trial with columns |
region_col |
Character name of the 0/1 region indicator column in
|
subgroup_index |
Optional logical vector of length |
pi_preserve |
Numeric threshold for the Method 1 consistency check:
the check passes when log(HR region) / log(HR overall) is at least
|
Value
A list with numeric vectors overall, non_region and region
(each with elements hr, lo, hi, n, d; NA HR and limits when
fewer than 10 subjects or 5 events), consistency_ratio, logical
method1_pass, method2_pass and region_hr_below_1, and, when
subgroup_index is supplied, region_subgroup and
non_region_subgroup.
Simulate Survival Data from AFT Data Generating Mechanism
Description
Generates simulated survival data from a previously created AFT data generating mechanism (DGM). Samples from the super population and generates survival times with specified censoring.
Usage
simulate_from_dgm(
dgm,
n = NULL,
rand_ratio = 1,
entry_var = NULL,
max_entry = 24,
analysis_time = 48,
cens_adjust = 0,
draw_treatment = TRUE,
seed = NULL,
strata_rand = NULL,
hrz_crit = NULL,
keep_rand = FALSE,
time_eos = NULL,
replace = TRUE,
baseline = c("resample", "fixed")
)
Arguments
dgm |
An object of class |
n |
Integer specifying the sample size. If |
rand_ratio |
Numeric randomisation ratio (treatment:control).
Default |
entry_var |
Character string naming an entry-time variable in the
super population. If |
max_entry |
Numeric maximum entry time for staggered entry simulation.
Only used when |
analysis_time |
Numeric calendar time of analysis. Follow-up is
|
cens_adjust |
Numeric log-scale adjustment to censoring distribution.
Positive values increase censoring times; negative values decrease them.
Default |
draw_treatment |
Logical. If |
seed |
Integer random seed. Default |
strata_rand |
Character string naming a column in the sampled data
for within-stratum balanced treatment allocation. If |
hrz_crit |
Numeric log-HR threshold. If supplied, a column
|
keep_rand |
Logical. If |
time_eos |
Numeric secondary administrative censoring cutoff
(end-of-study time on the DGM scale). Applied after |
replace |
Logical sampling scheme for drawing the |
baseline |
Character; how baseline covariates are obtained per
simulated trial. |
Details
Time-scale consistency
All time parameters (analysis_time, max_entry,
time_eos) must be expressed in the same units as
outcome_var supplied to generate_aft_dgm_flex(). A common
error is building the DGM on days (e.g. rfstime) and then passing
analysis_time in months, which causes follow-up windows far shorter
than the DGM event-time scale and produces universal administrative
censoring (event_sim = 0 for all subjects).
Verify with: exp(dgm$model_params$mu) — the implied median event
time should be plausible given your analysis_time.
n = NULL path
When n = NULL the entire super population is used as-is, with no
staggered entry and no administrative censoring (follow_up = Inf).
Treatment assignments and linear predictors already stored in
dgm$df_super are retained unchanged.
Censoring adjustment
cens_adjust shifts the log-scale location parameter of the
censoring distribution:
-
cens_adjust = log(2)doubles expected censoring times. -
cens_adjust = log(0.5)halves expected censoring times.
Random-X vs fixed-X designs (baseline)
Under baseline = "resample" each simulated trial is a fresh draw
of subjects, so between-trial variability reflects both covariate
resampling and outcome generation – the "what if this trial had enrolled
a different sample from the same population?" interpretation. Subgroup
sizes fluctuate across trials.
Under baseline = "fixed" the covariate matrix is the observed
trial data, identical in every simulated trial: subgroup definitions,
memberships, and sizes are exactly those of the source data. Between-trial
variability isolates the outcome-generation process (treatment
assignment, latent event times, censoring, entry). This is the
conditional-on-X (parametric-bootstrap-given-X) design. Note that entry
times are still drawn Uniform(0, max_entry) per trial unless
entry_var names a fixed column, and treatment is still
re-randomised per trial unless draw_treatment = FALSE.
Value
A data.frame with columns:
idSubject identifier.
treatOriginal treatment from super population.
treat_simSimulated treatment assignment.
flag_harmSubgroup indicator (1 = all subgroup conditions met).
z_*Covariate values.
lin_pred_1,lin_pred_0Counterfactual log-time linear predictors.
y_simObserved survival time (
min(T, C)).event_simEvent indicator (1 = event, 0 = censored).
t_trueLatent true survival time (pre-censoring).
c_timeEffective censoring time (post admin-censoring).
hrz_flag(Optional) Individual harm-zone indicator.
rand_order(Optional) Randomisation sequence index.
See Also
generate_aft_dgm_flex, check_censoring_dgm
Examples
## Not run:
dgm <- setup_gbsg_dgm(model = "null", verbose = FALSE)
sim_data <- simulate_from_dgm(dgm, n = 200, seed = 42)
dim(sim_data)
head(sim_data[, c("y_sim", "event_sim", "treat_sim")])
## End(Not run)
Simulate Trial Data from GBSG DGM
Description
Generates simulated clinical trial data from a GBSG-based data generating mechanism.
Usage
simulate_from_gbsg_dgm(
dgm,
n = NULL,
rand_ratio = 1,
sim_id = 1,
max_follow = Inf,
muC_adj = 0,
min_cens = NULL,
max_cens = NULL,
draw_treatment = TRUE
)
Arguments
dgm |
A "gbsg_dgm" object from |
n |
Integer. Sample size. If NULL, uses full super-population. Default: NULL |
rand_ratio |
Numeric. Randomization ratio (treatment:control). Default: 1 (1:1 randomization) |
sim_id |
Integer. Simulation ID used for seed offset. Default: 1 |
max_follow |
Numeric. Administrative censoring time (months). Default: Inf (no administrative censoring) |
muC_adj |
Numeric. Adjustment to censoring distribution location parameter. Positive values increase censoring. Default: 0 |
min_cens |
Numeric. Minimum censoring time for uniform censoring. Required if cens_type = "uniform" |
max_cens |
Numeric. Maximum censoring time for uniform censoring. Required if cens_type = "uniform" |
draw_treatment |
Logical. If TRUE, randomly assigns treatment. If FALSE, samples from existing treatment arms. Default: TRUE |
Value
Data frame with simulated trial data including:
- id
Subject identifier
- y.sim
Observed follow-up time
- event.sim
Event indicator (1 = event, 0 = censored)
- t.sim
True event time (before censoring)
- treat
Treatment indicator
- flag.harm
Harm subgroup indicator
- loghr_po
Individual log hazard ratio (potential outcome)
- v1-v7
Analysis factors
Examples
## Not run:
dgm <- create_gbsg_dgm(model = "alt", k_inter = 2.0)
sim_data <- simulate_from_gbsg_dgm(dgm, n = 500, sim_id = 1)
# Check AHR in simulated data
exp(mean(sim_data$loghr_po))
## End(Not run)
Simulate Trial Data from a GLM Data Generating Mechanism
Description
Generates a simulated clinical trial dataset by sampling from the
super-population in a "glm_dgm" object. Assigns treatment
and generates outcomes from the individual-level potential outcomes
stored in the DGM.
Usage
simulate_from_glm_dgm(
dgm,
n = NULL,
rand_ratio = 1,
seed = NULL,
draw_treatment = TRUE,
replace = TRUE,
baseline = c("resample", "fixed")
)
Arguments
dgm |
An object of class |
n |
Integer. Sample size. If |
rand_ratio |
Numeric. Treatment:control randomisation ratio.
Default |
seed |
Integer. Random seed. Default |
draw_treatment |
Logical. If |
replace |
Logical sampling scheme for drawing the |
baseline |
\"resample\" (default; draw the trial panel from the super-population) or \"fixed\" (use dgm$df_source – the observed panel, every patient exactly once, observed order; requires a DGM built by a version that stores df_source; n must be NULL or nrow(df_source)). |
Value
A data frame with columns:
idInteger subject identifier.
y_simSimulated outcome.
treat_simTreatment indicator (0/1).
flag_harmTrue subgroup membership indicator.
Plus all covariate columns from the super-population.
For binary outcomes, y_sim is a 0/1 Bernoulli draw from
the potential outcome probabilities (p0 or p1).
There is no event_sim column – for ForestSearch, set
event.name = "y_sim".
See Also
generate_glm_dgm,
run_simulation_analysis
Examples
## Not run:
dgm <- generate_glm_dgm(...)
df <- simulate_from_glm_dgm(dgm, n = 500, seed = 42)
table(df$y_sim, df$treat_sim)
## End(Not run)
Sort Subgroups by Focus (post-consistency)
Description
Sorts a data.table of subgroup results according to the specified focus.
For "hrMaxSG" and "hrMinSG", candidates are first
partitioned into a candidate inclusion set (the "band"), then
ranked within the band by sample size. The selection_rule
argument controls how the band is defined:
-
"neighborhood"– 1D effect-size band (default; legacy behaviour): a candidate is in-band iff its effect is withineffect_neighborhoodof the maximum. -
"pareto"– 2D Pareto-non-dominated set in (effect, N) space: a candidate is in-band iff no other candidate has both a larger effect and a larger N. -
"both"– intersection of the two: a candidate is in-band iff it satisfies the neighborhood test and is Pareto-non-dominated.
See Details.
Usage
sort_subgroups(
result_new,
sg_focus,
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
effect_log_scale = FALSE
)
Arguments
result_new |
A data.table of subgroup results with columns
|
sg_focus |
Character. Sorting focus. One of |
selection_rule |
Character. Rule defining the candidate
inclusion set for |
effect_neighborhood |
Numeric in |
effect_log_scale |
Logical. If |
Details
Sort keys by sg_focus:
"hr"(-Pcons, -hr, K)– prefer high consistency and large effect."maxSG"(-N, -Pcons, K)– prefer large subgroups with high consistency."minSG"(N, -Pcons, K)– prefer small subgroups with high consistency."hrMaxSG"(-in_band, -N, -Pcons, -hr, K)– among candidates in the inclusion band defined byselection_rule, prefer the largest sample size."hrMinSG"(-in_band, N, -Pcons, -hr, K)– among candidates in the inclusion band defined byselection_rule, prefer the smallest sample size.
"hrMaxSG" and "hrMinSG" use a lexicographic with
candidate band rule: effect size is primary, but a bounded set of
candidates near (or trading off against) the maximum is accepted to
optimize sample size.
Selection rules:
"neighborhood"1D
\varepsilon-band on the effect scale:in_band = (hr >= (1 - effect_neighborhood) * max(hr)). Settingeffect_neighborhood = 0reduces this to a strict max-effect filter. This is the legacy behaviour."pareto"2D Pareto-non-dominated set: a candidate is in-band iff no other candidate has both
hr_j >= hr_iandN_j >= N_iwith at least one strict inequality. Reuses the same dominance computation ascompute_pareto_frontier(). Removes theeffect_neighborhoodtuning parameter from the selection criterion."both"Intersection: in-band iff in the
\varepsilon-band and Pareto-non-dominated. Strictest of the three.
Value
A sorted data.table. The top row is the selected subgroup
under sg_focus; remaining rows are diagnostic.
Examples
library(data.table)
dt <- data.table(Pcons = c(0.92, 0.95, 0.88, 0.90),
hr = c(2.5, 2.4, 2.3, 2.0),
N = c(70, 100, 150, 200),
K = c(1, 2, 1, 2))
# hrMaxSG with 10% neighborhood: HR 2.5, 2.4, 2.3 are in-band;
# among those, N = 150 is largest -> top row is N=150, hr=2.3.
sort_subgroups(dt, sg_focus = "hrMaxSG", effect_neighborhood = 0.10)
# Same data with selection_rule = "pareto": the dominance set in
# (hr, N) is used instead of the 1D effect band.
sort_subgroups(dt, sg_focus = "hrMaxSG", selection_rule = "pareto")
Sort Subgroups by Focus (pre-consistency)
Description
Sorts a data.table of candidate subgroups before consistency
evaluation. Mirrors sort_subgroups but operates on the
pre-consistency column convention (HR, n, K)
and omits Pcons tiebreakers (consistency not yet available).
Usage
sort_subgroups_preview(
result_new,
sg_focus,
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
effect_log_scale = FALSE
)
Arguments
result_new |
A data.table of candidate subgroups with columns
|
sg_focus |
Character. Sorting focus. One of |
selection_rule |
Character. See |
effect_neighborhood |
Numeric in |
effect_log_scale |
Logical. See |
Value
A sorted data.table.
Evaluate Subgroup Consistency
Description
Evaluates candidate subgroups using split-sample consistency validation. For each candidate, repeatedly splits the data and checks whether the treatment effect direction is consistent across splits.
Usage
subgroup.consistency(
df,
hr.subgroups,
hr.threshold = 1,
hr.consistency = 1,
pconsistency.threshold = 0.9,
m1.threshold = Inf,
n.splits = 100,
details = FALSE,
by.risk = 12,
plot.sg = FALSE,
maxk = 7,
Lsg,
confs_labels,
sg_focus = "hr",
selection_rule = "neighborhood",
effect_neighborhood = 0.1,
stop_Kgroups = 200,
stop_threshold = NULL,
show_candidate_summary = FALSE,
max_print = 10,
pconsistency.digits = 2,
seed = 8316951,
checking = FALSE,
use_twostage = FALSE,
twostage_args = list(),
parallel_args = list(),
estimator_fn = NULL,
consistency_threshold = NULL,
effect_label = "HR",
effect_log_scale = FALSE,
adjust_covariates = NULL,
consistency_method = c("resample", "split"),
glm_resample_spec = NULL
)
Arguments
df |
Data frame containing the analysis dataset. Must include columns for outcome (Y), event indicator (Event), and treatment (Treat). |
hr.subgroups |
Data.table of candidate subgroups from subgroup search, containing columns: HR, n, E, K, d0, d1, m0, m1, grp, and factor indicators. |
hr.threshold |
Numeric. |
hr.consistency |
Numeric. |
pconsistency.threshold |
Numeric. |
m1.threshold |
Numeric. Maximum m1 threshold for filtering. Default: Inf |
n.splits |
Integer. Number of splits for consistency evaluation. Default: 100 |
details |
Logical. Print progress details. Default: FALSE |
by.risk |
Numeric. Risk interval for KM plots. Default: 12 |
plot.sg |
Logical. Generate subgroup plots. Default: FALSE |
maxk |
Integer. Maximum number of factors in subgroup. Default: 7 |
Lsg |
List of subgroup parameters. |
confs_labels |
Character vector mapping factor names to labels. |
sg_focus |
Character. Subgroup selection criterion. One of:
Default: |
selection_rule |
Character. Rule defining the candidate
inclusion set for |
effect_neighborhood |
Numeric in |
stop_Kgroups |
Integer. Maximum number of candidates to evaluate.
When the ranked pool exceeds this cap it is truncated and a warning naming
both counts and the |
stop_threshold |
Numeric in Note: Values > 1.0 are not permitted. To disable early
stopping, use Interaction with
For parallel execution, early stopping is checked after each batch
completes, so some additional candidates beyond the first meeting the
threshold may be evaluated. Use a smaller |
show_candidate_summary |
Logical. If |
max_print |
Integer. Maximum number of candidate rows printed in
each |
pconsistency.digits |
Integer. Number of decimal places to which the
consistency proportion is rounded before it is compared with
|
seed |
Integer. Random seed for reproducible consistency splits. Default: 8316951. Set to NULL for non-reproducible random splits. The seed is used both for sequential execution (via set.seed()) and parallel execution (via future.seed). |
checking |
Logical. Enable additional validation checks. Default: FALSE |
use_twostage |
Logical. Use two-stage adaptive algorithm. Default: FALSE |
twostage_args |
List. Parameters for two-stage algorithm:
|
parallel_args |
List. Parallel processing configuration:
|
estimator_fn |
Closure or |
consistency_threshold |
Numeric or |
effect_label |
Character. Column label for the effect measure in
diagnostic output (e.g., candidate subgroup table). Default |
effect_log_scale |
Logical. If |
adjust_covariates |
Character vector or |
consistency_method |
Character. |
glm_resample_spec |
List or |
Value
A list containing:
- out_sg
Selected subgroup results. When non-
NULL, containsresult(sorted candidate table; top row is the selected subgroup) andpareto_frontier(data.table of non-dominated candidates on (effect, N), both maximized – a post-hoc diagnostic, not used for selection).- sg_focus
Selection criterion used
- df_flag
Data frame with treatment recommendations
- sg.harm
Subgroup definition labels – character vector of factor-level cut names (e.g.,
c("z1.1", "z2.1")), of length equal to the number of cuts defining the subgroup.NULLif no subgroup was identified.- sg.harm.id
Per-subject subgroup-membership indicator – integer vector of length
nrow(df)with1if subjectiis in the identified subgroup and0otherwise.NULLif no subgroup was identified. Not a character vector of cut expressions; see the Field naming collision section below.- algorithm
"twostage" or "fixed"
- n_candidates_evaluated
Number of candidates actually evaluated
- n_candidates_total
Total candidates available
- n_passed
Number meeting consistency threshold
- early_stop_triggered
Logical indicating if early stop occurred
- early_stop_candidate
Index of candidate triggering early stop
- stop_threshold
Threshold used for early stopping
- seed
Random seed used for reproducibility (NULL if not set)
Field naming collision with GRF results
The field name sg.harm.id has different semantics on
this object versus on the result objects returned by
grf.subg.harm.glm and grf.subg.harm.survival:
| Object | sg.harm.id contains | Length / type |
subgroup.consistency() result (this function) | per-subject 0/1 membership indicator | integer, length nrow(df) |
grf.subg.harm.glm() result | character vector of cut expressions | character, length = depth of selected tree |
grf.subg.harm.survival() result | character vector of cut expressions | character, length = depth of selected tree |
Practical consequence. Do not paste sg.harm.id with
paste(..., collapse = " & ") to print "the identified subgroup"
without first checking object class. For subgroup labels from a
forestsearch or subgroup.consistency() result,
use sg.harm (character vector of cut names) – the FS main
result exposes sg.harm at the top level.
This naming collision is a documented CRAN-stable API for v0.1.x and v0.2.x; it is expected to be resolved via a deprecation cycle in a future minor release.
Examples
## Not run:
# Standard evaluation
result <- subgroup.consistency(
df = trial_data,
hr.subgroups = candidates,
sg_focus = "hr",
n.splits = 400,
parallel_args = list(plan = "multisession", workers = 6)
)
# Show top 10 candidates before evaluation
result <- subgroup.consistency(
df = trial_data,
hr.subgroups = candidates,
sg_focus = "hr",
show_candidate_summary = TRUE, # Post-consistency summary
n.splits = 400
)
# With early stopping and custom batch size
result <- subgroup.consistency(
df = trial_data,
hr.subgroups = candidates,
sg_focus = "hr",
stop_threshold = 0.95,
show_candidate_summary = TRUE,
parallel_args = list(
plan = "multisession",
workers = 6,
batch_size = 2 # Check early stopping after every 2 candidates
)
)
## End(Not run)
Subgroup Search for Treatment Effect Heterogeneity (Improved, Parallelized)
Description
Searches for subgroups with treatment effect heterogeneity using combinations of candidate factors. Evaluates subgroups for minimum prevalence, event counts, and hazard ratio threshold. Parallelizes the main search loop.
Usage
subgroup.search(
Y,
Event,
Treat,
ID = NULL,
Z,
n.min = 30,
d0.min = 15,
d1.min = 15,
hr.threshold = 1,
max.minutes = 30,
minp = 0.05,
rmin = 5,
details = FALSE,
maxk = 2,
parallel_workers = parallel::detectCores(),
estimator_fn = NULL,
df_analysis = NULL,
effect_threshold = NULL,
effect_measure = NULL,
outcome_type = NULL,
adjust_covariates = NULL,
disable_effect_floor = FALSE
)
Arguments
Y |
Numeric vector of outcome (e.g., time-to-event). |
Event |
Numeric vector of event indicators (0/1). |
Treat |
Numeric vector of treatment group indicators (0/1). |
ID |
Optional vector of subject IDs. |
Z |
Matrix or data frame of candidate subgroup factors (binary indicators). |
n.min |
Integer. Minimum subgroup size. |
d0.min |
Integer. Minimum per-arm filter. For survival and binary
outcomes: minimum number of events in the control arm within the
candidate subgroup. Ignored for continuous outcomes (only
|
d1.min |
Integer. Same as |
hr.threshold |
Numeric. |
max.minutes |
Numeric. Currently inert; scheduled for
deprecation in v0.3.0. Previously intended as a wall-clock budget for
the combination search, it is threaded to
|
minp |
Numeric. Minimum prevalence rate for each factor. |
rmin |
Integer. Minimum required reduction in sample size when adding a factor. |
details |
Logical. Print details during execution. |
maxk |
Integer. Maximum number of factors in a subgroup. |
parallel_workers |
Integer. Number of parallel workers (default: all available cores). |
estimator_fn |
Closure or |
df_analysis |
Data frame or |
effect_threshold |
Numeric or |
effect_measure |
Character or |
outcome_type |
Character or |
adjust_covariates |
Character vector or |
disable_effect_floor |
Logical. When |
Value
List with found subgroups, maximum HR, search time, configuration info, and filtering statistics.
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
Z <- df[, c("age", "meno", "size", "grade3", "nodes", "pgr", "er")]
res <- subgroup.search(Y = df$rfstime, Event = df$status,
Treat = df$hormon, Z = Z,
hr.threshold = 1.25)
## End(Not run)
Cox per-subgroup fit function
Description
Returns a function data -> c(HR, UB) fitting formula by
survival::coxph() and extracting the first row of
summary(fit)$conf.int: the hazard-ratio point estimate
("exp(coef)") and the upper 95% bound ("upper .95"). Estimation
failure or an empty coefficient table returns c(NA, NA). This
reproduces the vignettes' cox_sR() exactly when called with
survival::Surv(y_sim, event_sim) ~ treat_sim + survival::strata(grade).
Usage
subgroup_cox(formula, lean = TRUE)
Arguments
formula |
A model formula for |
lean |
Re-home the formula environment as described above
(default |
Details
Custom fitters may be passed to run_subgroup_sims() directly: any
function of a data frame returning a length-2 numeric
(estimate, upper bound) is accepted. Fitters run on parallel
workers, so they should reference non-base functions by
pkg::fun or rely on packages listed in future_packages.
Value
A function of one argument (data) returning
c(hr, ub); carries the formula in attr(, "formula").
Formula environment (lean)
A formula captures the environment where it was created. In a script
or a knitr chunk that is the global environment, and future's
globals machinery gives formulas special treatment: a fit whose
formula points at a populated environment imposes a multi-second
per-future cost on parallel runs (measured ~2-4 s x ~116 futures in
the Phase 2 timing diagnostic), independent of the serialized size.
With lean = TRUE (default) the formula is re-homed into a fresh,
empty environment parented on the survival namespace, cutting the
costly chain while keeping survival::Surv(), survival::strata(),
and base/stats symbols resolvable. Consequence: the formula must
reference only data columns and functions reachable from the
survival namespace (use pkg:: prefixes for anything else). Set
lean = FALSE to retain the calling environment – restoring
arbitrary-symbol resolution at that parallel cost. Independently of
lean, the returned closure is stripped of source references and
given a minimal environment holding only the formula (keep-source
installs otherwise attach ~250 KB of package-source baggage to every
worker shipment).
Examples
fit <- subgroup_cox(survival::Surv(y_sim, event_sim) ~ treat_sim)
GLM per-subgroup fit function
Description
Returns a function data -> c(estimate, UB) for non-survival outcomes
by delegating estimation to the package's native effect-estimator
layer: the closure from make_effect_estimator() (the same code the
forestsearch() search, bootstrap, and consistency loops use),
including the bit-identical unadjusted fast path
(.make_lm_estimator_fast()) under the same routing predicate
forestsearch() applies. The estimator's internal-scale
estimate/se are mapped to the wrapper contract as
upper = estimate + qnorm(1 - (1 - level)/2) * se, back-transformed
(exp()) for ratio measures – the package's own display convention
– so ratio-measure matrices land on the natural scale exactly as
subgroup_cox()'s exp(coef) / upper .95 do. Estimation failure,
non-convergence, NA estimate/SE (e.g. a single-arm subgroup, where
the treatment coefficient is aliased), or a non-finite result (e.g. a
zero-residual-df fit) returns c(NA, NA).
Usage
subgroup_glm(
outcome.name = "y_sim",
treat.name = "treat_sim",
outcome_type = "continuous",
effect_measure = NULL,
adverse_outcome = TRUE,
robust_se = TRUE,
offset.name = NULL,
adjust_covariates = NULL,
level = 0.95,
lean = TRUE
)
Arguments
outcome.name, treat.name |
Column names of the outcome and the
0/1 treatment indicator; defaults match |
outcome_type |
|
effect_measure |
|
adverse_outcome |
Higher outcome = harm ( |
robust_se, offset.name, adjust_covariates |
Passed to
|
level |
Confidence level for the upper bound (default |
lean |
Re-home the provenance formula's environment (default
|
Value
A function of one argument (data) returning
c(estimate, ub); carries the provenance formula in
attr(, "formula") and scale metadata in attr(, "effect") – a
list with measure, outcome_type, log_scale, null_value,
adverse_outcome, est_label, ub_label, est_thresholds,
ub_thresholds, level – which run_subgroup_sims() stamps onto
GLM results for scale-aware summaries and plots. A custom fitter
can opt into the same behavior by attaching an "effect" attribute
of this shape.
Direction convention
Estimates follow the estimator layer's harm-positive normalization:
with adverse_outcome = FALSE (higher outcome = better) the outcome
is negated internally so that estimate > 0 consistently indicates
treatment harm, aligning with the survival convention where HR > 1 =
harm. Consequently the null value is always 0 (identity measures) or
1 (ratio measures), whatever the raw outcome's direction.
Supported settings
Phase 4.5 validates outcome_type = "continuous" (MD),
outcome_type = "binary" ("OR" – the binary default – and
"RD"), and outcome_type = "count" ("IRR" – the count default
– and "IRD"; offset.name is required, see below). Binary
"RR" (estimator-supported; wrapper parity pending) still stops
with an informative error; the estimator layer
(make_effect_estimator()) already supports it and the gate lifts
with its parity phase – no interface change.
Serialization (lean)
The returned closure is stripped of source references and given a
minimal environment (the estimator closure and two scalars), and the
estimator closure itself is rebuilt onto a fresh environment holding
only its forced argument values (parented on the package namespace,
which serializes by reference) with its own source references
stripped. This severs shared srcfile environments (keep-source
installs) and the constructor's frame chain from the serialized
graph, so the fitter ships to parallel workers in the low
single-digit KB regardless of session state. The provenance formula
attached in attr(, "formula") (built from the column names, e.g.
y_sim ~ treat_sim) is re-homed into an empty stats-parented
environment when lean = TRUE, keeping the parallel-shipment guard
in run_subgroup_sims() meaningful.
See Also
subgroup_cox(), run_subgroup_sims(),
make_effect_estimator(), generate_glm_dgm()
Examples
## Not run:
fit <- subgroup_glm() # y_sim ~ treat_sim, MD
fit(simulate_from_glm_dgm(dgm, n = 500, seed = 1))
# Binary (Phase 4.4): OR by default, RD available
fit_or <- subgroup_glm(outcome_type = "binary")
fit_rd <- subgroup_glm(outcome_type = "binary", effect_measure = "RD")
# Count (Phase 4.5): IRR by default, IRD available; offset required
fit_irr <- subgroup_glm(outcome_type = "count", offset.name = "t_exp")
fit_ird <- subgroup_glm(outcome_type = "count", offset.name = "t_exp",
effect_measure = "IRD")
## End(Not run)
Suggest Screening and Consistency Thresholds
Description
Recommends (c1, c2) threshold pairs for ForestSearch based on
the corrected null approximation. Uses the per-subgroup detection
probability and an L_{\text{eff}} multiplicity correction to
find the smallest c1 (screening threshold) such that the
corrected FPR is at most fpr_target.
Usage
suggest_thresholds(
d_eff,
N,
fpr_target = 0.1,
n_min = 60,
L_eff_C = 1,
L_eff_alpha = 0,
c1_range = c(1.05, 2.5),
c1_step = 0.05,
c2_candidates = seq(0.9, 2, by = 0.05),
effect_scale = c("ratio", "difference")
)
Arguments
d_eff |
Numeric. Effective information in the expected harm
subgroup. Use |
N |
Integer. Total sample size. |
fpr_target |
Numeric. Maximum acceptable procedure-level FPR. Default: 0.10. |
n_min |
Integer. Minimum subgroup size. Default: 60. |
L_eff_C |
Numeric. Calibration constant for
|
L_eff_alpha |
Numeric. Calibration exponent. Default: 0.0 (no N-dependence; see Details). |
c1_range |
Numeric vector of length 2. Search range for c1.
Default: |
c1_step |
Numeric. Grid step for c1. Default: 0.05. |
c2_candidates |
Numeric vector. Candidate c2 values to evaluate
for each c1. Default: |
effect_scale |
Character. |
Details
The corrected FPR is:
\text{FPR}_{\text{corr}} = 1 - (1 - P_1)^{L_{\text{eff}}}
where P_1 = compute_detection_probability_glm(1.0, d_eff, c1, c2)
and L_{\text{eff}} = C (N / n_{\min})^\alpha.
Default L_eff values (C=1, alpha=0) assume L_{\text{eff}} = 1,
which gives the per-subgroup approximation with no multiplicity
correction. This is appropriate when:
No calibration data is available
The analysis has few confounders and small N
For more accurate suggestions, calibrate L_{\text{eff}} from
an H0 simulation using calibrate_L_eff or
run_null_calibration.
Reference calibrations (from the forestsearch GLM extension analysis documents):
- Binary, 4 confounders
C = 0.220, alpha = 1.298 (binary_threshold_calibration.qmd, 5000 sims)
- GBSG survival, 7 confounders
C = 0.029, alpha = 0.882 (cox_vs_glm_approximation.qmd, 500 sims per N)
Value
A data.frame with columns:
- c1
Screening threshold.
- min_c2
Minimum consistency threshold achieving the target.
- FPR_corr
Corrected FPR at this (c1, min_c2).
- P1
Per-subgroup detection probability.
- L_eff
Effective number of candidates at this N.
Sorted by c1 ascending. The first row is the smallest c1 that achieves the target FPR.
See Also
compute_detection_probability_glm,
calibrate_L_eff,
predict_fpr_corrected
Examples
# Survival: paper defaults should be recovered
d <- d_eff_survival(n_sg = 60, prop_cens = 0.45)
suggest_thresholds(d_eff = d, N = 700)
# Binary: needs higher thresholds
d <- d_eff_binary(n_sg = 150, p_event = 0.30)
suggest_thresholds(d_eff = d, N = 500,
L_eff_C = 0.220, L_eff_alpha = 1.298)
# Use calibration object from run_null_calibration()
# cal <- run_null_calibration(...)
# suggest_thresholds(d_eff = d, N = 700,
# L_eff_C = cal$C, L_eff_alpha = cal$alpha)
Summarize Bootstrap Event Counts
Description
Provides summary statistics for event counts across bootstrap iterations, helping assess the reliability of HR estimates when events are sparse.
Usage
summarize_bootstrap_events(boot_results, threshold = 5)
Arguments
boot_results |
List. Output from forestsearch_bootstrap_dofuture() |
threshold |
Integer. Minimum event threshold for flagging low counts (default: 5) |
Details
This function summarizes event counts in four scenarios:
ORIGINAL subgroup H evaluated on BOOTSTRAP samples
ORIGINAL subgroup Hc evaluated on BOOTSTRAP samples
NEW subgroup H* (found in bootstrap) evaluated on ORIGINAL data
NEW subgroup Hc* (found in bootstrap) evaluated on ORIGINAL data
Low event counts (below threshold) can lead to unstable HR estimates. This summary helps identify potential issues with sparse events.
Value
Invisibly returns a list with summary statistics:
- threshold
The event threshold used
- nb_boots
Total number of bootstrap iterations
- n_successful
Number of iterations that found a new subgroup
- original_H
List with low event counts for original H on bootstrap samples
- original_Hc
List with low event counts for original Hc on bootstrap samples
- new_Hstar
List with low event counts for new H* on original data
- new_Hcstar
List with low event counts for new Hc* on original data
Examples
## Not run:
# After running bootstrap analysis
summarize_bootstrap_events(fs_bc, threshold = 10)
## End(Not run)
Enhanced Bootstrap Results Summary
Description
Creates comprehensive output including formatted table with subgroup footnote, diagnostic plots, bootstrap quality metrics, and detailed timing analysis.
Usage
summarize_bootstrap_results(
sgharm,
boot_results,
create_plots = FALSE,
est.scale = "hr",
digits = 2
)
Arguments
sgharm |
The selected subgroup object from forestsearch results. Can be:
|
boot_results |
List. Output from forestsearch_bootstrap_dofuture() |
create_plots |
Logical. Generate diagnostic plots (default: FALSE) |
est.scale |
Character. "hr" or "1/hr" for effect scale |
digits |
Integer or NULL. Decimal places for display formatting of
numeric columns in the rendered table (per-arm summaries, Diff,
effect estimate CI, bias-corrected CI). Default 2. When set, the
formatted strings in |
Details
The table output includes a footnote displaying the identified subgroup
definition, analogous to the tab_estimates table from sg_tables.
This is achieved by extracting the subgroup definition from sgharm and
passing it to format_bootstrap_table.
Value
List with components:
- table
gt table with treatment effects and subgroup footnote
- diagnostics
List of bootstrap quality metrics
- diagnostics_table_gt
gt table of diagnostics
- plots
List of ggplot2 diagnostic plots (if create_plots=TRUE)
- timing
List of timing analysis (if timing data available)
- subgroup_summary
List from summarize_bootstrap_subgroups()
See Also
format_bootstrap_table for table creation
sg_tables for analogous main analysis tables
summarize_bootstrap_subgroups for subgroup stability analysis
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
fs <- forestsearch(df,
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"),
outcome.name = "rfstime", treat.name = "hormon", event.name = "status")
fs_bc <- forestsearch_bootstrap_dofuture(fs, nb_boots = 100)
summaries <- summarize_bootstrap_results(fs$sg.harm, fs_bc)
summaries$table
## End(Not run)
Summarize Bootstrap Subgroup Analysis Results
Description
Comprehensive tabulation of bootstrap subgroup identification results
including basic statistics, factor frequencies, consistency
distributions, GRF-cut frequencies, and agreement with the original
analysis subgroup. Returns raw data.table tabulations
suitable for custom post-processing, export, or plotting; for the
full formatted pipeline (gt tables, diagnostics, plots) call
summarize_bootstrap_results instead.
Usage
summarize_bootstrap_subgroups(results, nb_boots, original_sg = NULL, maxk = 2)
Arguments
results |
Data.table or data.frame. Bootstrap results with subgroup characteristics including columns like Pcons, hr_sg, N_sg, K_sg, and M.1-M.k |
nb_boots |
Integer. Total number of bootstrap iterations |
original_sg |
Character vector. Original subgroup definition from main analysis (e.g., c("{age>=50}", "{nodes>=3}") for a 2-factor subgroup) |
maxk |
Integer. Maximum number of factors allowed in subgroup definition |
Value
List with summary components:
- basic_stats
Data.table of summary statistics
- consistency_dist
Data.table of Pcons distribution by bins
- size_dist
Data.table of subgroup size distribution
- factor_freq
Data.table of factor frequencies by position
- agreement
Data.table of subgroup definition agreement counts
- factor_presence
Data.table of base factor presence counts
- factor_presence_specific
Data.table of specific factor definitions
- grf_cut_freq
Data.table of GRF policy-tree cut frequencies across all bootstrap iterations (not just successful ones). Columns:
Rank,grf_cut,N,Percent.NULLwhenresultslacks thegrf_cuts_bcolumn (pre-Phase-C objects). Unlike subgroup-related tables, this tabulation is NOT restricted to iterations that identified a subgroup – it spans all bootstraps so that the analyst can distinguish "GRF produced no cut" from "GRF produced a cut that consistency rejected".- original_agreement
Data.table comparing to original analysis subgroup
- n_found
Integer. Number of successful iterations
- pct_found
Numeric. Percentage of successful iterations
See Also
summarize_bootstrap_results for the full
formatted summary pipeline (gt tables plus diagnostics and plots).
forestsearch_bootstrap_dofuture to run the bootstrap
analysis that produces the results object consumed here.
Examples
## Not run:
# Run bootstrap analysis
boot_out <- forestsearch_bootstrap_dofuture(fs.est, nb_boots = 500)
# Raw tabulations for custom post-processing
subg <- summarize_bootstrap_subgroups(
results = boot_out$results,
nb_boots = 500,
original_sg = fs.est$sg.harm,
maxk = 2L
)
# Inspect the GRF cut frequencies across all bootstraps
subg$grf_cut_freq
# Export factor frequencies to CSV for an external report
data.table::fwrite(subg$factor_freq, "factor_freq.csv")
## End(Not run)
Summarise ForestSearch Cross-Validation Diagnostics
Description
Post-hoc diagnostic summary of forestsearch_tenfold output,
aggregating the sim x fold grid captured in fold_summary into
publication-ready gt tables. Parallel to
summarize_bootstrap_results but focused on cross-validation
stability and GRF candidate-discovery behaviour.
Usage
summarize_cv_results(
cv_output,
original_sg = NULL,
original_grf_cuts = NULL,
create_plots = FALSE,
top_n = 15L
)
Arguments
cv_output |
List of class |
original_sg |
Character vector or |
original_grf_cuts |
Character vector or |
create_plots |
Logical. If |
top_n |
Integer. Maximum number of detail rows to retain in
frequency tables (default: |
Value
An object of class "fs_cv_summary", a list with
components:
identification_summary,grf_cut_summary,cut_vs_subgroup_xtab,no_subgroup_decomposition,pconsistency_distribution,fold_numeric_summary,original_agreement,metrics_tablegttables (ordata.framefallback if gt is unavailable). Some may beNULLdepending on input columns and optional arguments.plotsNULLor a list ofggplotobjects.dataList of raw
data.frametabulations underlying the tables (identification,grf_cuts,cut_vs_subgroup,no_subgroup,pconsistency,fold_numeric_summary,original_agreement).n_sims,n_folds,total_pairsGrid dimensions.
has_grf_cuts,has_pconsistencyLogical. Whether the corresponding
fold_summarycolumns were available. Thefold_numeric_summaryslot adapts automatically to whichever numeric diagnostic columns are present and does not have a dedicated flag.
Diagnostic components
The returned object includes (as gt tables by default):
-
identification_summary: top identified subgroups across all sim x fold pairs, with counts and percentages. -
grf_cut_summary: top rawgrf_cutsstrings across training sets – how often GRF returned each output (and how often it returned none).NULLiffold_summary$grf_cutsis not available (e.g., objects from forestsearch prior to the GLM-extension branch). -
cut_vs_subgroup_xtab: cross-tabulation of uniquegrf_cutsstrings against identified subgroups. Reveals whether GRF's discovery aligns with the pipeline's final selection. -
no_subgroup_decomposition: tabulation of folds that identified no subgroup, broken into "GRF returned no cut" vs "GRF returned a cut but consistency rejected it". -
pconsistency_distribution: binned distribution of the consistency probability (Pcons) achieved by the identified subgroup on each fold, restricted to folds that identified a subgroup; includes median and IQR. Reveals whether identifying folds are scrapingpconsistency.threshold(e.g., Pcons just above 0.90) or are well above it.NULLiffold_summary$pconsistencyis not available. Rejected- candidate Pcons values are not surfaced – only the selected candidate's Pcons is captured infold_summary. -
fold_numeric_summary: tidy median / IQR / range table with one row per numeric diagnostic column present infold_summary(n_test,pconsistency,training_fs_hr,n_candidates_evaluated). Absent columns simply don't produce rows, so the table adapts to pre-feature-branch objects without error. Training-fold subgroup effects are on the effect measure's natural scale (HR for survival; OR, RR, IRR, RD, IRD, or MD for GLM, exponentiated at capture for ratio measures so the column is always on natural scale regardless of outcome type); they are in-sample estimates (optimistically biased) and are surfaced for diagnostic comparison across folds only. -
original_agreement: iforiginal_sgand/ororiginal_grf_cutsare supplied, fraction of folds matching the original full-data analysis (exact subgroup, partial shared factor, GRF-cut match, both). -
metrics_table: formatted version ofcv_output$sens_summaryandcv_output$find_summary(delegates tocv_metrics_tables). -
plots: ifcreate_plots = TRUE, bar charts of top identified subgroups and top GRF cuts.
Raw data.frame versions of each tabulation are available at
$data$* for custom post-processing.
Backward compatibility
The GRF-dependent slots (grf_cut_summary,
cut_vs_subgroup_xtab, no_subgroup_decomposition) require
a grf_cuts column in fold_summary; the
pconsistency_distribution slot requires a pconsistency
column. Both were added on the feature/glm-extension branch
(grf_cuts first, pconsistency later). Older objects
are detected and the affected slots return NULL without error.
Metadata fields has_grf_cuts and has_pconsistency
record which columns were present.
See Also
forestsearch_tenfold for running repeated
K-fold CV; cv_metrics_tables for the underlying
metrics formatter; summarize_bootstrap_results for
the bootstrap analogue.
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
df$time_months <- df$rfstime / 30.4375
fs <- forestsearch(
df.analysis = df,
confounders.name = c("age", "meno", "size", "grade3",
"nodes", "pgr", "er"),
outcome.name = "time_months",
event.name = "status",
treat.name = "hormon"
)
tf <- forestsearch_tenfold(fs.est = fs, sims = 100, Kfolds = 10)
# Basic diagnostic summary
cv_diag <- summarize_cv_results(tf)
cv_diag$identification_summary
cv_diag$grf_cut_summary
cv_diag$no_subgroup_decomposition
# With original-analysis comparison and plots
cv_diag_full <- summarize_cv_results(
tf,
original_sg = fs$sg.harm,
original_grf_cuts = fs$grf_cuts,
create_plots = TRUE
)
cv_diag_full$original_agreement
cv_diag_full$plots$identification
## End(Not run)
Summarize Factor Presence Across Bootstrap Subgroups
Description
Analyzes how often each individual factor appears in identified subgroups, extracting base factor names from full definitions and identifying common specific definitions.
Usage
summarize_factor_presence_robust(
results,
maxk = 2,
threshold = 10,
as_gt = TRUE
)
Arguments
results |
Data.table or data.frame. Bootstrap results with M.1, M.2, etc. columns |
maxk |
Integer. Maximum number of factors allowed |
threshold |
Numeric. Percentage threshold for including specific definitions (default: 10) |
as_gt |
Logical. Return gt tables (TRUE) or data.frames (FALSE) |
Value
List with base_factors and specific_factors data.frames or gt tables
Summarize Simulation Results
Description
Creates a summary table of operating characteristics across all simulations. Includes both HR and AHR metrics.
Usage
summarize_simulation_results(
results,
analyses = NULL,
digits = 2,
digits_hr = 3
)
Arguments
results |
data.table with simulation results from run_simulation_analysis |
analyses |
Character vector. Analysis methods to include. Default: all |
digits |
Integer. Decimal places for proportions. Default: 2 |
digits_hr |
Integer. Decimal places for hazard ratios. Default: 3 |
Value
Data frame with summary statistics
Examples
## Not run:
dgm <- create_gbsg_dgm()
sim_res <- run_simulation_analysis(dgm, nsim = 10)
summarize_simulation_results(sim_res)
## End(Not run)
Summarize Single Analysis Results
Description
Summarize Single Analysis Results
Usage
summarize_single_analysis(result, digits = 2, digits_hr = 3)
Arguments
result |
data.table with results for a single analysis method |
digits |
Integer. Decimal places for proportions |
digits_hr |
Integer. Decimal places for hazard ratios |
Value
Data frame with summary statistics
Summary method for cox_ahr_cde objects
Description
Summary method for cox_ahr_cde objects
Usage
## S3 method for class 'cox_ahr_cde'
summary(object, ...)
Arguments
object |
A |
... |
Additional arguments (not used). |
Value
Invisibly returns the input object.
Examples
## Not run:
library(survival)
df <- survival::gbsg
df$grade3 <- as.integer(df$grade == "3")
res <- cox_ahr_cde_analysis(df,
outcome.name = "rfstime", event.name = "status", treat.name = "hormon",
confounders.name = c("age", "meno", "size", "grade3", "nodes", "pgr", "er"))
summary(res)
## End(Not run)
Summary method for a DINA fit.
Description
Builds a coefficient table with estimate, standard error, Wald z-statistic, and two-sided p-value for each coefficient.
Usage
## S3 method for class 'dina'
summary(object, ...)
Arguments
object |
an object of class |
... |
unused. |
Value
an object of class "summary.dina", printable via
print().
Summary Method for forestsearch Objects
Description
Provides a detailed summary of a ForestSearch analysis including input parameters, variable selection results, consistency evaluation, and the selected subgroup with key metrics.
Usage
## S3 method for class 'forestsearch'
summary(object, ...)
Arguments
object |
A |
... |
Additional arguments (currently unused). |
Value
Invisibly returns object.
Post-selection inference
Carries the same block as print.forestsearch() (see there) and adds the
log-scale standard errors (field, field-s with the naive complement SE
beside it, and IJ two-term), the three largest re-selection frequencies
with their labels, and p_hat_sum over the re-selection family. Absent MR
results, the output is unchanged.
When the fit carries the membership-agreement diagnostics – that is, under
mr_inference_args = list(field_recovery = TRUE) – one further line
reports sens_H, the mean share of the identified subgroup that the field's
re-selections retain, beside \hat p. The pair separates two very
different low-\hat p situations: \hat p counts exact
re-selections of \hat H, so a low \hat p with a high sens_H
says the draws re-selected near-twins of \hat H, while a low
\hat p with a low sens_H says they re-selected unrelated regions.
Like \hat p it is descriptive and no construction reads it, and it
answers a narrower question than a bootstrap or cross-validation recovery
rate: re-selection within the fixed candidate family under perturbation, not
re-discovery from resampled data. The line is absent – and the rest of the
output byte-identical – when the diagnostics were not requested.
print.forestsearch() does not report it.
Examples
## Not run:
fs <- forestsearch(df.analysis = mydata, ...)
summary(fs)
## End(Not run)
Summary method for fpr_calibration objects
Description
Delegates to print.fpr_calibration.
Usage
## S3 method for class 'fpr_calibration'
summary(object, ...)
Arguments
object |
An |
... |
Passed to the print method. |
Value
The input object, invisibly.
Summarize an extreme-subgroups simulation study
Description
Computes, from the raw matrices of a run_subgroup_sims() result (or
a loaded RDS payload coerced with class<-), every statistic used by
the vignettes' results table and forest panels: 1st/50th/99th
percentile matrices for the estimate and its upper bound, convergence
counts, mean subgroup N, tail probabilities, unconditional median
display strings, the formatted results_tbl, category labels and
colours, validity masks, the single/combination panel index vectors,
and the high-risk panel (threshold filter, ITT anchor, ordering).
Usage
## S3 method for class 'subgroup_sims'
summary(
object,
hr_true = object$hr_true,
probs = c(0.01, 0.5, 0.99),
est_thresholds = NULL,
ub_thresholds = NULL,
est_label = NULL,
ub_label = NULL,
single_cats = c("ITT", "Clinical", "Continuous"),
combo_cats = c("Interaction", "3-way", "Random"),
pr_thresh_highrisk = 0.1,
itt_name = "All Patients",
cat_cols = c(ITT = "black", Clinical = "steelblue", Continuous = "darkgreen",
Interaction = "purple4", `3-way` = "violetred3", Random = "darkorange"),
...
)
Arguments
object |
A |
hr_true |
True effect on the fitter's estimate scale, for
reference lines (a hazard ratio for Cox fits, a mean difference
for |
probs |
Quantile probabilities for the ECI matrices; the default
|
est_thresholds |
Length-2 numeric |
ub_thresholds |
Length-2 numeric: the upper-bound tails are
|
est_label, ub_label |
Display labels for the estimate and its
upper bound in the table headers (legacy |
single_cats, combo_cats |
Category labels routed to the single-variable and combination forest panels. |
pr_thresh_highrisk |
High-risk panel inclusion threshold on
the first upper-bound tail probability, default |
itt_name |
Name of the ITT row, always included in and anchored at the top of the high-risk panel. |
cat_cols |
Named colour map for categories. |
... |
Unused. |
Details
Field names are structural and retained across outcome types
(hr_q, pr_hr_lt050, pr_ub_ge2, ...), exactly as sim_hrs and
hr.threshold serve generic duty elsewhere in the package: on a GLM
result they hold the fitter's estimate-scale statistics. Estimates
follow the estimator layer's harm-positive normalization, so the
upper tail (Pr(est > null), Pr(UB >= t)) always points toward
harm whatever the raw outcome's direction.
Threshold and label resolution: explicit arguments win; otherwise the
result's effect metadata supplies them (subgroup_glm() stamps
scale-aware defaults from Phase 4.4 – ratio measures carry the
legacy tails c(0.5, 1) / c(2, 3), identity measures
c(NA, null_value) / c(NA, NA) – with labels from the effect
measure); otherwise the HR legacy values
c(0.5, 1.0) / c(2, 3) with the vignettes' literal headers. An
NA threshold degrades gracefully: its probability field is NA
and its table column renders "-" under a placeholder header.
Structurally empty subgroups (all-NA columns) yield NA medians
and NaN means/proportions, exactly as in the vignettes; the
formatted table renders them as "-".
Value
An object of class "subgroup_sims_summary": a list with
fields named as in the vignettes (hr_q, ub_q, n_valid,
pct_valid, mean_n, pr_hr_lt050, pr_hr_gt1, pr_ub_ge2,
pr_ub_ge3, mhr_uncond, mub_uncond, results_tbl,
sg_names, cat, cat_cols, ok, ok_ub, idx_hr_single,
idx_hr_combo, idx_ub_single, idx_ub_combo, n_single,
n_combo, highrisk, n_sims, design, hr_true). Legacy
calls return exactly this set; effect-aware calls append effect
(the result's metadata, possibly NULL when only explicit
overrides were given), thresholds (list(est =, ub =) as
resolved), and labels (list(est =, ub =)).
See Also
run_subgroup_sims(), subgroup_glm()
Summary Tables for MRCT Simulation Results
Description
Creates summary tables from MRCT simulation results using the gt package. Summarizes hazard ratio estimates, subgroup identification rates, and classification of identified subgroups. Optionally displays two scenarios (e.g., alternative and null hypotheses) side by side.
Usage
summaryout_mrct(
pop_summary = NULL,
mrct_sims,
mrct_sims_null = NULL,
scenario_labels = c("Alternative", "Null"),
pop_summary_null = NULL,
sg_type = 1,
tab_caption = "Identified subgroups and estimation summaries",
digits = 3,
trim_threshold = 1000,
trim_fraction = 0.01,
table_width = 600,
font_size = 11,
showtable = TRUE
)
Arguments
pop_summary |
List. Population summary from large sample approximation (optional). Default: NULL |
mrct_sims |
data.table. Simulation results from
|
mrct_sims_null |
data.table. Optional second set of simulation results (e.g., null hypothesis). When supplied, the table displays two value columns side by side. Default: NULL (single-scenario table). |
scenario_labels |
Character vector of length 2. Column headers for the
two scenarios. Only used when |
pop_summary_null |
List. Population summary for the null scenario (optional). Default: NULL |
sg_type |
Integer. Type of subgroup summary: 1 = basic summary (found, biomarker, age); 2 = extended summary (all subgroup types). Default: 1 |
tab_caption |
Character. Caption for the output table. Default: "Identified subgroups and estimation summaries" |
digits |
Integer. Number of decimal places for numeric summaries. Default: 3 |
trim_threshold |
Numeric. When the raw mean of a metric exceeds this
value in absolute terms, the summary switches to a symmetrically trimmed
mean and SD (excluding the lower and upper |
trim_fraction |
Numeric between 0 and 0.5. Fraction of observations to trim from each tail when trimming is triggered. Default: 0.01 (1 percent from each tail, i.e., the central 98 percent of values). |
table_width |
Numeric. Total table width in pixels. Column widths are allocated proportionally. Increase for HTML/wide displays (e.g., 750), decrease for beamer slides (e.g., 550). Default: 600. |
font_size |
Numeric. Base font size in pixels. Title is
|
showtable |
Logical. Print the table. Default: TRUE |
Value
List with components:
- res
List of summary statistics from population. When dual-scenario, contains
res_altandres_null.- out_table
Formatted gt table object, or data.frame if gt is unavailable.
- data
Processed mrct_sims data.table with derived variables. When dual-scenario, also contains
data_null.- summary_df
Data frame of computed summary statistics.
See Also
mrct_region_sims for generating simulation results
Examples
## Not run:
# Single scenario (backward-compatible)
summaryout_mrct(
mrct_sims = results_alt,
tab_caption = "H1: Heterogeneous treatment effect"
)
# Dual scenario: alternative vs null side-by-side
summaryout_mrct(
mrct_sims = results_alt,
mrct_sims_null = results_null,
scenario_labels = c("Alternative (HTE)", "Null (uniform)"),
tab_caption = "Operating characteristics: Alternative vs Null"
)
# Custom trimming: 2% from each tail when mean > 500
summaryout_mrct(
mrct_sims = results_alt,
mrct_sims_null = results_null,
trim_threshold = 500,
trim_fraction = 0.02
)
## End(Not run)
Test for Constant Conditional Average Treatment Effect
Description
H0': tau(x) = tau for some tau and all x.
Usage
test_constant_cate(
Y,
W,
X,
poly_order = 1L,
covariate_select = c("all", "top_down", "bottom_up"),
t_threshold = 2,
regression = c("ols", "logistic", "poisson"),
offset = NULL
)
Arguments
Y |
Numeric outcome vector. |
W |
Treatment indicator (0/1). |
X |
Numeric covariate matrix (no factors; use
|
poly_order |
Integer. Default: 1. |
covariate_select |
"all", "top_down", or "bottom_up". |
t_threshold |
Numeric. Default: 2.0. |
regression |
"ols", "logistic", or "poisson". |
offset |
Numeric offset vector (Poisson only). |
Value
A list with components test, chi_sq, df,
p_value_chi, normal (normal-approximation statistic),
p_value_normal, K (number of covariates), N0,
N1, regression, covariates_selected,
ate_diff (intercept-difference point estimate), and
diff_slope, V_slope. Returns NULL if per-arm
fits fail or fewer than two covariates are retained.
Run Both Crump et al. Tests
Description
Run Both Crump et al. Tests
Usage
test_hte(
Y,
W,
X,
poly_order = 1L,
covariate_select = c("all", "top_down", "bottom_up"),
t_threshold = 2,
regression = c("ols", "logistic", "poisson"),
offset = NULL
)
Arguments
Y |
Numeric outcome vector. |
W |
Treatment indicator (0/1). |
X |
Numeric covariate matrix (no factors; use
|
poly_order |
Integer. Default: 1. |
covariate_select |
"all", "top_down", or "bottom_up". |
t_threshold |
Numeric. Default: 2.0. |
regression |
"ols", "logistic", or "poisson". |
offset |
Numeric offset vector (Poisson only). |
Value
Object of class hte_test.
HTE Test Using ForestSearch-Selected Covariates
Description
HTE Test Using ForestSearch-Selected Covariates
Usage
test_hte_from_forestsearch(
fs_result,
df,
outcome.name,
treat.name,
regression = c("ols", "logistic", "poisson"),
offset = NULL,
poly_order = 1L
)
Arguments
fs_result |
A |
df |
Data frame used in the |
outcome.name |
Character. Outcome variable name. |
treat.name |
Character. Treatment variable name. |
regression |
Character: "ols", "logistic", or "poisson". |
offset |
Numeric offset vector (Poisson only). |
poly_order |
Integer. Default: 1. |
Value
List with hte_test, covariates_used, source.
Test for Zero Conditional Average Treatment Effect
Description
H0: tau(x) = 0 for all x.
Usage
test_zero_cate(
Y,
W,
X,
poly_order = 1L,
covariate_select = c("all", "top_down", "bottom_up"),
t_threshold = 2,
regression = c("ols", "logistic", "poisson"),
offset = NULL
)
Arguments
Y |
Numeric outcome vector. |
W |
Treatment indicator (0/1). |
X |
Numeric covariate matrix (no factors; use
|
poly_order |
Integer. Default: 1. |
covariate_select |
"all", "top_down", or "bottom_up". |
t_threshold |
Numeric. Default: 2.0. |
regression |
"ols", "logistic", or "poisson". |
offset |
Numeric offset vector (Poisson only). |
Value
List with test results including covariates_selected.
Tidy a cut / subgroup expression for display
Description
Display-only cleanup for diagnostic printing of cut and subgroup-definition
strings. Collapses runs of whitespace to single spaces (so padded GRF
definitions like "cd40 <= 320" render cleanly) and
rounds each post-operator numeric threshold half-up to digits places
(matching the collapse_cuts representative rounding, e.g.
"wtkg > 79.833600000000004" -> "wtkg > 80" at digits = 0).
Usage
tidy_cut_display(s, digits = 0L)
Arguments
s |
Character vector of expressions, e.g.
|
digits |
Integer >= 0. Rounding for embedded thresholds. Default 0
(nearest integer), matching the |
Details
Only numbers that follow a comparison operator (<=, >=,
==, <, >) are rounded, so integers embedded in variable
names (cd40, prior_6mo) are never altered. This function does
NOT touch the canonical, membership-exact labels used in the search; it is
intended only for the diagnostic strings passed to cat().
Value
Character vector, tidied for display (same length as s).
Validate input data for GRF analysis
Description
Checks that input data meets requirements for GRF analysis
Usage
validate_grf_data(W, D, n.min)
Arguments
W |
Numeric vector. Treatment indicator |
D |
Numeric vector. Event indicator |
n.min |
Integer. Minimum subgroup size |
Value
Logical. TRUE if data is valid, FALSE with warning otherwise
Examples
W <- rep(0:1, each = 50)
D <- rbinom(100, 1, 0.6)
validate_grf_data(W, D, n.min = 60)
Validate Input Parameters
Description
Validate Input Parameters
Usage
validate_inputs(
data,
model,
cens_type,
outcome_var,
event_var,
treatment_var,
continuous_vars,
factor_vars
)
Validate k_inter Effect on HR Heterogeneity
Description
Test function to verify that k_inter properly modulates the difference between HR(H) and HR(Hc), and that AHR metrics align with Cox-based HRs.
Usage
validate_k_inter_effect(
k_inter_values = c(-2, -1, 0, 1, 2, 3),
verbose = TRUE,
...
)
Arguments
k_inter_values |
Numeric vector of k_inter values to test. Default: c(-2, -1, 0, 1, 2, 3) |
verbose |
Logical. Print results. Default: TRUE |
... |
Additional arguments passed to create_gbsg_dgm |
Value
Data frame with k_inter, hr_H, hr_Hc, AHR_H, AHR_Hc, and ratio columns
Examples
## Not run:
# Test k_inter effect
results <- validate_k_inter_effect()
# k_inter = 0 should give hr_H approximately equals hr_Hc (ratio approximately 1)
## End(Not run)
Validate Dataset for MRCT Simulations
Description
Checks that a dataset contains all required variables for MRCT simulation functions and reports any issues. Required variables include outcome (tte, event), treatment (treat), continuous covariates (age, bm), and factor covariates (male, histology, prior_treat, regA).
Usage
validate_mrct_data(df.case, verbose = TRUE)
Arguments
df.case |
Data frame to validate |
verbose |
Logical. Print detailed validation results. Default: TRUE |
Details
Required Variables
The function checks for the following variables:
-
Outcome: tte (time-to-event), event (0/1 indicator)
-
Treatment: treat (0/1 indicator)
-
Continuous: age, bm (biomarker)
-
Factor: male (0/1), histology, prior_treat (0/1), regA (0/1)
The function also validates variable types and value ranges.
Value
Logical. TRUE if all requirements met, FALSE otherwise (invisibly)
See Also
create_dgm_for_mrct for creating DGM from validated data
Examples
## Not run:
# Check if dataset is ready for MRCT simulations
is_valid <- validate_mrct_data(df.case)
if (!is_valid) {
stop("Please fix data issues before running simulations")
}
## End(Not run)
Validate Spline Specification
Description
Validate Spline Specification
Usage
validate_spline_spec(spline_spec, df_work)
Validate subgroup definitions against a data frame
Description
Evaluates every subgroup id expression against data (augmented
with flag_itt = 1L, the scalar cutpoints as columns, and 0-valued
benchmark membership columns) and fails with a named, collected error
if any expression errors, is non-logical, or has the wrong length.
Run automatically at the top of run_subgroup_sims() so that a typo
in a subgroup definition stops the study in milliseconds instead of
surfacing as a silent all-NA column after the simulation loop.
Usage
validate_subgroups(
subgroups,
data,
cutpoints = list(),
benchmarks = benchmark_spec()
)
Arguments
subgroups |
List of subgroup definitions; each element a list
with character scalars |
data |
A data frame carrying the covariates the |
cutpoints |
Named list of scalars exposed as columns (e.g.
|
benchmarks |
A |
Value
Invisibly TRUE; errors otherwise.
Wilson Score Confidence Interval
Description
Computes Wilson score confidence interval for a proportion, which has better coverage properties than the normal approximation for small samples and proportions near 0 or 1.
Usage
wilson_ci(x, n, conf.level = 0.95)
Arguments
x |
Integer. Number of successes. |
n |
Integer. Number of trials. |
conf.level |
Numeric. Confidence level (default 0.95). |
Value
Named numeric vector with elements: estimate, lower, upper.
Examples
wilson_ci(90, 100)
wilson_ci(5, 20, conf.level = 0.90)