aiEvalR: Statistical and Psychometric Evaluation of AI Systems

Evaluates artificial intelligence (AI) systems as measurement instruments using psychometric methods. Provides multi-facet generalizability theory (G-study and D-study) via 'lme4', reliability via the intraclass correlation coefficient (ICC), calibration via the expected calibration error (ECE) and Brier score, robustness stress testing, and group disparity diagnostics. Item-level differential item functioning (DIF) based on item response theory (IRT) is delegated to the 'aiDIF' package. Methods follow Cronbach, Gleser, Nanda and Rajaratnam (1972, <ISBN:9780471188506>) and Brennan (2001) <doi:10.1007/978-1-4757-3456-0>.

Version: 0.1.0
Depends: R (≥ 4.1.0)
Imports: stats, utils
Suggests: testthat (≥ 3.0.0), knitr, rmarkdown, covr, boot, lme4, dplyr, ggplot2, aiDIF, spelling
Published: 2026-08-30
DOI: 10.32614/CRAN.package.aiEvalR (may not be active yet)
Author: Subir Hait ORCID iD [aut, cre]
Maintainer: Subir Hait <haitsubi at msu.edu>
BugReports: https://github.com/causalfragility-lab/aiEvalR/issues
License: MIT + file LICENSE
URL: https://github.com/causalfragility-lab/aiEvalR
NeedsCompilation: no
Language: en-US
Materials: README, NEWS
CRAN checks: aiEvalR results

Documentation:

Reference manual: aiEvalR.html , aiEvalR.pdf
Vignettes: Getting Started with aiEvalR (source, R code)
The Generalizability-Theory Core (source, R code)

Downloads:

Package source: aiEvalR_0.1.0.tar.gz
Windows binaries: r-devel: not available, r-release: not available, r-oldrel: not available
macOS binaries: r-release (arm64): aiEvalR_0.1.0.tgz, r-oldrel (arm64): aiEvalR_0.1.0.tgz, r-release (x86_64): aiEvalR_0.1.0.tgz, r-oldrel (x86_64): aiEvalR_0.1.0.tgz

Linking:

Please use the canonical form https://CRAN.R-project.org/package=aiEvalR to link to this page.