Operating manual
Tag: evaluation
A focused reading path for evaluation: related field notes, evidence trails, and operating questions from the archive.
- voice-agents-need-a-different-reliability-test.md
Voice Agents Need a Different Reliability Test
Voice agent reliability should track intent from capture through action. This six-stage operator framework finds where meaning first breaks.
open artifact → - your-agent-is-learning-the-harness-too.md
Your AI Agent Is Learning the Harness, Too
When training runs through an execution harness, the interface shapes learned behavior. Use this six-part contract to test what will transfer.
open artifact →