Operating manual
Tag: Agent Operations
A focused reading path for Agent Operations: related field notes, evidence trails, and operating questions from the archive.
- scoring-definition-is-the-lever-in-agent-evals.md
Scoring Definition Is the Lever in Agent Evals
A 24.0%-57.0% consensus swing on identical held-out data is a property of the scoring definition, not the agents. Lock the definition before reading the number.
open artifact → - agent-depth-trades-yield-for-cheaper-context.md
Agent Depth Trades Yield for Cheaper Context
Depth is not automatically cheaper context. Treating finding yield, root-context exposure, and dollars per finding as three independent axes exposes the real trade, and a per-tier alignment check prevents promoting another tier on paper numbers alone.
open artifact → - benchmark-agent-decisions-not-just-final-scores.md
Benchmark Agent Decisions, Not Just Final Scores
A final score hides whether an agent used evidence well. Evaluate the decision trace: context, intervention, consequence, validity, and budget.
open artifact → - the-last-wrong-step-is-not-where-the-agent-failure-started.md
The Last Wrong Step Is Not Where Failure Started
A four-stage operator model for tracing long-horizon agent errors from their initiating deviation through propagation, resolution, and terminal impact.
open artifact → - your-ai-security-checklist-has-a-version-problem.md
Your AI Security Checklist Has a Version Problem
A practical control for pinning AI security guidance, preserving its source, and reopening reviews when that guidance changes.
open artifact → - voice-agents-need-a-different-reliability-test.md
Voice Agents Need a Different Reliability Test
Voice agent reliability should track intent from capture through action. This six-stage operator framework finds where meaning first breaks.
open artifact →