benchmarks
Prompt Monitoring Without Misleading Yourself
A defensible monitoring program documents what a run observed, what changed, and what the sample cannot establish.
Practical research on AI discovery
Research section
Methods, metrics, and limits for measuring how information appears in AI answers.
2 articles
benchmarks
A defensible monitoring program documents what a run observed, what changed, and what the sample cannot establish.
benchmarks
Seven bounded measures for a documented prompt sample, with formulas, failure modes, and hypothetical examples—not claims of universal engine performance.