Operating manual
Tag: evaluation
A focused reading path for evaluation: related field notes, evidence trails, and operating questions from the archive.
- build-the-workflow-before-the-prompt.md
Build the Workflow Before the Prompt
Build verified tool-use examples from executable workflows. What ToolGrad shows, what its pass rate does not, and how to keep challenge cases separate.
open artifact → - ai-benchmarks-must-name-the-serving-route.md
AI Benchmarks Must Name the Serving Route
Before transferring an AI benchmark score to your deployment, record the serving configuration, capability coverage, and failures behind it.
open artifact → - a-model-upgrade-is-a-memory-migration.md
A Model Upgrade Is a Memory Migration
Copied records do not prove an agent kept its memory. Model upgrades need directional tests, clean index rebuilds, repair evidence, and rollback.
open artifact → - quantum-error-mitigation-break-even-point.md
Quantum Error Mitigation Has a Break-Even Point
Quantum error mitigation must earn its overhead under a fixed budget. Compare squared bias and variance to classify it as helped, hurt, or indistinguishable.
open artifact → - benchmark-agent-decisions-not-just-final-scores.md
Benchmark Agent Decisions, Not Just Final Scores
A final score hides whether an agent used evidence well. Evaluate the decision trace: context, intervention, consequence, validity, and budget.
open artifact → - a-world-model-is-an-executable-hypothesis.md
A World Model Is an Executable Hypothesis
A test-time world model should earn authority through replay, bounded action, and immediate revocation when reality produces a counterexample.
open artifact →