NVIDIA's AVO harness ran Claude Opus 5 autonomously for seven days straight
NVIDIA paired Claude Opus 5 with a new harness called AVO to run long-horizon autonomous tasks, including a seven-day GPU kernel optimisation run and a separate reasoning benchmark. AVO uses persistent memory and a supervisor process so the agent keeps working past a single context window instead of restarting.
Why it matters
Most agent demos run for minutes. Getting an agent to make real progress over days, across many restarts and context resets, needs more than a longer context window: something has to persist what was tried, and something has to notice when the agent is stuck. NVIDIA's result suggests this pattern generalises across unrelated task types rather than being a one-off trick.
What changed
NVIDIA's technical blog described AVO, an architecture for long-horizon autonomous agents, tested by pairing Claude Opus 5 with the harness on two long-running tasks: a seven-day GPU kernel optimisation run, and a separate attempt at the ARC-AGI-3 reasoning benchmark. AVO relies on two mechanisms. Persistent memory carries forward prior implementations, evaluation results, compiler and profiler output, and accumulated reasoning, so the agent resumes from its current state instead of reconstructing the search. A supervisor process monitors the broader trajectory for stagnation or repeated unproductive cycles and can redirect the agent toward alternative strategies, while the agent itself still decides what to inspect, change, test and evaluate. NVIDIA said AVO performed well on both task types, suggesting it generalises rather than being tuned to one benchmark.
In an interview
A candidate could describe the persistent-memory-plus-supervisor pattern as a practical answer to "how do you keep an agent working on something for days": externalise state so progress survives a context reset, and add a separate process whose only job is noticing when the main loop is stuck.
Just read this? Test yourself on what you have been reading.
Read the original → martinfowler.com
- #ai-agents
- #nvidia
- #claude-opus-5
- #agent-architecture
- #long-horizon-tasks
You just looked this up. Could you explain it out loud?
That is the part interviews actually test. Tezvyn takes questions like this one and gives you what the interviewer is really checking, the answer that lands, and the mistake that ends the conversation, in the four minutes before your next meeting.
The iPhone app is on the way
We are building it. Until it lands, nothing here is held back from you: every interview card, your saved cards, streaks and the job board all work in Safari, plus hundreds of free practice quizzes of thirty questions each. Sign in and it all carries over to the app the day it arrives.
Want it as an icon? Tap Share at the bottom of Safari, then Add to Home Screen. It opens full screen and the cards you have read stay available offline.
We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.
See open roles