---
title: "Process-IRT Model Atlas: What to Fit, What to Validate, What Not to Claim"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Process-IRT Model Atlas: What to Fit, What to Validate, What Not to Claim}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(eyeprocess)
```

# Purpose

The process-IRT layer is deliberately organized by *measurement question*, not
by estimator novelty. Eye-tracking, pupillometry, response time, omissions, and
sequences become explicit measurement channels only when their role and
validation evidence are stated.

```{r}
validation_evidence_levels()
list_irt_models()
```

# Core model families

| Question | Primary API | Default scientific status |
|---|---|---|
| Do response, time, and gaze share person/item structure? | `fit_joint_gaze_rt_irt()` | reference/experimental |
| Do graded scores and time/process co-vary? | `fit_joint_graded_rt_process_irt()` | experimental |
| Which option was chosen and inspected? | `fit_nominal_gaze_irt()` | reference/experimental |
| Does visual exposure inform missingness? | `fit_gaze_informed_missingness_irt()` | diagnostic |
| Are omissions and not-reached items time processes? | `fit_omission_survival_irt()` | reference/experimental |
| Are process measures transportable across device/session/algorithm? | `fit_manyfacet_process_irt()` | reference |
| Does the response process change within a session? | `fit_changepoint_multimodal_irt()` | experimental |
| Do latent sequence states relate to measurement? | `fit_process_hmm_irt()` | experimental |
| Do process features explain DIF nuisance variation? | `audit_process_adjusted_dif()` | diagnostic |
| Is there residual person-item geometry? | `fit_latent_space_irt()` | external engine |
| Are logistic IRFs too restrictive? | `fit_gpirt()` | model criticism/gated |
| Does a bounded process outcome pile up at 0/1? | `fit_censored_normal_process_irt()` | conditional calibration |
| Are event times informative conditional on theta? | `fit_event_time_irt()` | diagnostic/gated |
| Are multiple selected options informative beyond a total score? | `fit_multiple_response_process_irt()` | reference/external gated |
| Is there residual inter-option/process dependence? | `audit_process_local_dependence()` | diagnostic |
| Do revisits/RT/gaze add evidence to cognitive diagnosis? | `fit_revisit_process_cdm()` | adapter/experimental |
| Does a process channel add held-out information? | `audit_channel_incremental_information()` | validation |

# A process channel must earn its place

The preferred comparison is not “model with gaze has a lower in-sample AIC.”
Instead, compare held-out performance and run a negative control.

```{r, eval=FALSE}
inc <- audit_channel_incremental_information(
  data = trials,
  fold = "participant_id",
  baseline_fitter = fit_without_gaze,
  process_fitter = fit_with_gaze,
  predictor = predict_model,
  scorer = score_model,
  higher_is_better = TRUE
)
plot(inc)

neg <- negative_control_process_test(
  data = trials,
  process = "dwell_time",
  fold = "participant_id",
  fitter = fit_with_gaze,
  predictor = predict_model,
  scorer = score_model
)
plot(neg)
```

# Missingness: separate exposure from response

```{r, eval=FALSE}
miss <- classify_item_missingness(
  trials,
  response = "response",
  reached = "reached",
  inspected = "inspected",
  started = "response_started"
)

fit <- fit_gaze_informed_missingness_irt(
  trials,
  response = "response",
  person = "participant_id",
  item = "item_id",
  gaze_exposure = "item_dwell_ms",
  theta = "theta"
)
plot(fit)
```

A fitted association between gaze exposure and omission is not evidence that
missingness is ignorable, nor is it a behavioral diagnosis. The two-part
reference model is intended to expose this dependency before a fully joint
missingness model is claimed.

# Cross-device measurement is an estimand

```{r, eval=FALSE}
facets <- fit_manyfacet_process_irt(
  trials,
  response = "correct",
  process = "dwell_ms",
  person = "participant_id",
  item = "item_id",
  device = "device",
  session = "session",
  algorithm = "fixation_algorithm"
)

device_facet_effects(facets, channel = "process")
session_facet_effects(facets, channel = "process")
algorithm_facet_effects(facets, channel = "process")
audit_process_measurement_invariance(facets)
```

A small device variance component is not enough for interchangeability. It
should be accompanied by semantic round-trip evidence, unit/coordinate audits,
and held-device/session validation.

# Latent distribution and IRF stress tests

```{r, eval=FALSE}
audit_latent_distribution(theta)
compare_latent_distribution_models(theta)
latent_distribution_stress_test(validation_runner)

shape <- fit_gpirt(response_matrix, engine = "spline_reference")
plot_irf_uncertainty(shape, item = 1)
cmp <- compare_parametric_nonparametric_irf(response_matrix, shape)
audit_irf_shape(cmp)
```

The spline-reference route is intentionally called a shape audit, not GPIRT.
Exact GPIRT, dynamic GPIRT, flow-MIRT, variational IRT, and full continuous-time
IRT remain behind explicit external-engine gates until validated implementations
are supplied.

# Promotion is evidence-based

```{r, eval=FALSE}
spec <- irt_validation_spec("joint_gaze_rt", replications = 500)

# retained recovery/SBC/PPC/transport results are combined into an evidence bundle
grade_model_evidence(evidence_bundle)
```

At minimum retain recovery, bias/RMSE, interval coverage, convergence/failure
classification, misspecification stress tests, preprocessing sensitivity, and
grouped/external validation. Bayesian models additionally require SBC and
posterior predictive checks; posterior SBC is appropriate when calibration near
the observed-data regime matters and the model-specific self-consistency
contract has been implemented.
