AI now writes the code, screens the candidates, and drafts the
decisions. We build the instruments that measure how these systems
actually behave, and keep the evidence that holds up when it counts.
Language models carry latent decision dispositions: identical scenarios
with one attribute swapped can produce different judgments, inconsistently
across vendors and shifting silently between versions. In regulated
contexts (insurance claims and underwriting, candidate screening,
lending) those dispositions carry legal consequence, and no validated
way to measure them exists.
Our work is early-stage measurement science: procedurally generated
matched-pair instruments with established reliability and validity
bounds, calibration against implanted dispositions in open
models, drift detection for stochastic, version-unstable
systems, and contamination resistance for public evaluation
grammars. The deliverable is not a score; it is a methodology whose
results survive adversarial scrutiny, and the evidence file that goes
with it.
Principal Investigator
Terry Harmer has spent more than twenty years building
measurement systems inside regulated enterprises, most recently as a
continuous improvement executive at a Fortune 250 insurer. MSc,
Management & Information Systems, Trinity College Dublin. His career
specialty is the discipline this research formalizes: measurement whose
results hold up in front of executives, auditors, and regulators.