Operating manual
Tag: ai-evaluation
A focused reading path for ai-evaluation: related field notes, evidence trails, and operating questions from the archive.
- agent-audit-budgets-need-a-control-group.md
Agent Audit Budgets Need a Control Group
Test confidence-ranked human review against independent baselines, then promote it only when it reduces measured risk at the same audit budget.
open artifact → - the-last-wrong-step-is-not-where-the-agent-failure-started.md
The Last Wrong Step Is Not Where Failure Started
A four-stage operator model for tracing long-horizon agent errors from their initiating deviation through propagation, resolution, and terminal impact.
open artifact → - confidence-should-not-allocate-your-audit-budget.md
Confidence Should Not Allocate Your Audit Budget
Agent confidence is one audit signal, not the allocator. Use a five-factor scorecard and random reserve when human review capacity is scarce.
open artifact → - judge-ai-by-the-decisions-it-improves.md
Judge AI by the Decisions It Improves
Judgment leverage is a five-part scorecard for testing whether AI improves accountable human decisions—not just output speed.
open artifact →