Nvidia research finds agentic harness design, not model capability, drives performance on long-horizon tasks
A custom harness with memory management and a supervisor component lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3, outperforming rival setups that lacked these features.
1 source · cross-referenced
- Nvidia research argues that the software harness around an AI model is more critical than the model itself for long-horizon agentic tasks.
- A custom harness with memory management and a supervisor component enabled Claude Opus 5 to reach 100% on the ARC-AGI-3 benchmark, versus 30% without the harness.
- OpenAI separately reported that tweaking two harness settings tripled its models’ scores on ARC-AGI-3, but none reached 100%.
- Nvidia’s Agentic Variation Operators (AVO) harness is not a product but part of its open Nemo toolkit for building agent systems.
Nvidia researchers report that the software "harness" around an AI model—tools, memory management, and rules—matters more than the model itself for long-horizon agentic tasks. In experiments, a custom harness with strong memory handling and a "supervisor" component enabled Claude Opus 5 to achieve a 100% score on the interactive reasoning benchmark ARC-AGI-3, a set of 2D games where the model must learn to play and win without instructions.
Without the harness, Opus 5 scored 30%, which was the top result among all models tested. The harness added a supervisor role that nudges the agent when it strays or risks dead ends, acting "almost like a CEO," according to Adel El Hallak, Nvidia’s vice president of product in its AI unit.
OpenAI separately disclosed that tweaking two harness settings tripled its models’ scores on ARC-AGI-3, though none reached 100%. Nvidia’s harness, called Agentic Variation Operators (AVO), is not a product but part of its open Nemo toolkit for building agent systems, including commercial and openly available components.
Nvidia argues that open harnesses give users more control over accuracy and costs than relying solely on proprietary models. The company contrasts this with OpenAI’s reported slowdown in model training due to security incidents, advocating for an open agent stack spanning harness, infrastructure, and runtime to advance the ecosystem securely.
- Aug 22, 2026 · Latent Space — swyx
Simile raises $2B Series B to build digital twins of human behavior for enterprise simulation
Trust71 - Aug 21, 2026 · Latent Space — swyx
Poolside strikes non-exclusive licensing deal with Nvidia, founders retain control with $1B pool
Trust71 - Aug 21, 2026 · Latent Space — swyx
New /wayfinder skill aims to reduce planning friction for autonomous agents
Trust78