How to evaluate harnesses for long-running autonomous agents
Uses the AVO case to show why teams evaluating long-running autonomous agents should prioritize state persistence, failure recovery, execution loops, and supervision over single-call model performance.























