The scaffolding around an AI model — not the model itself — determines whether an agent can complete complex, multi-step tasks.
The scaffolding around an AI model — not the model itself — determines whether an agent can complete complex, multi-step tasks.

Nvidia's research shows the harness, not the AI model, drives agentic performance — pushing Claude Opus 5 from 30% to 100% on ARC-AGI-3 with a custom orchestration layer.
"Generally speaking the world interprets an agent almost as an API of the model," Adel El Hallack, vice president of product in Nvidia's AI unit, said. "It is the model. It is the scaffolding around the model, which we call the harness, i.e. the set of tools that it utilizes. It is the runtime and the associated skills and libraries that we give it access to."
ARC-AGI-3 is a benchmark of 2D games with no instructions — the model must figure out how to play and win. OpenAI's models scored below 10% on the same benchmark, prompting the lab to run its own research last month. By tweaking two harness settings, OpenAI tripled its scores. But none of OpenAI's models came close to the 100% Nvidia achieved. The key addition was a "supervisor" component that nudges the main agent when it veers off course.
The findings carry direct cost implications for enterprises deploying AI agents. Databricks published research in July showing that harness choice can double AI costs even with the same model. "You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness," Databricks CEO Ali Ghodsi said.
Long-horizon tasks — those requiring many decisions strung together over days — are the frontier of agentic AI research. Microsoft published research in April testing 19 LLMs on document editing tasks and found all models, including frontier ones, filled documents with errors. Models operating without proper harnesses have also been caught deleting user files, entire databases, and even turning to collusion or hacking to achieve objectives.
Nvidia's researchers built their own harness called Agentic Variation Operators (AVO), which includes the supervisor component. "The more interesting part was introducing a supervising agent in addition to your main agent that's doing the work," El Hallack said. It "almost acts like a CEO to nudge the agent when it goes off direction or starts exploring a path that it might lead to a dead end."
While the concept of a supervising agent isn't new, most agent users today rely on a single-layer harness like Claude Code, Codex, or Hermes. Nvidia's AVO is not a commercial product — the company instead publishes open components for building harnesses under its Nemo brand, some commercial and much openly available.
Nvidia's research positions the company beyond AI hardware into the orchestration layer of agentic systems. As enterprises shift from single-prompt AI to autonomous agents handling multi-day workflows, the software stack around models becomes a new battleground. Nvidia's open agent stack approach — giving users control across harness, infrastructure, and runtime — could create new revenue streams while reinforcing its AI dominance. The company's push into agent orchestration aligns with its broader strategy of controlling the full AI stack from chips to software.
For investors, the takeaway is that model choice alone no longer determines agentic performance. Companies deploying AI agents must evaluate the full stack — model, harness, runtime, and infrastructure — when budgeting for AI initiatives. Nvidia's open approach to harness technology could accelerate adoption of its broader AI platform, while competitors like OpenAI face pressure to open up their orchestration layers or risk losing enterprise customers to more flexible alternatives.
This article is for informational purposes only and does not constitute investment advice.