Skip to content

What the software around a model is called, and what it measurably does.

The same language model can do far better or far worse work depending on the software around it. The measurements are large, and they come with limits that matter if you want to rely on them.

Ahsan Fazal

Founder and CEO

Note

The name

The software that lets a model act as an agent has had two names in three years. In July 2023 the evaluation group ARC Evals, now METR, wrote: “We call such programs ‘scaffolding’ and call scaffolding + model combinations ‘agents.’” By 2026 the word had become harness. Anthropic defines “an agent harness (or scaffold)” in its article on evaluating agents of 9 January 2026, as the system that “processes inputs, orchestrates tool calls, and returns results”. OpenAI wrote about harness engineering on 11 February 2026. Three preprints of May and June 2026 define it formally; one calls it “the infrastructure layer that governs context construction, tool interaction, orchestration, and verification around a language model” (Zhang and others, 7 May 2026).

We use harness. We looked for the word rig as a general term for this layer and found it in none of the sources above. Where it appears in AI tooling it is a proper name: a Rust library for LLM agents, and a security company that launched on 29 September 2026. In one agent-engineering source, Steve Yegge’s Gas Town, a rig is a project, one git repository under orchestration, and sits inside the harness. That is absence of evidence, not proof: informal use may exist where we did not look.

What the harness moves

Anthropic wrote in October 2024 that the performance of an agent on SWE-⁠bench “can vary significantly based on this scaffolding, even when using the same underlying AI model”. Later work puts numbers on it. Harness-Bench, of 27 May 2026, ran six harnesses on eight model backends over the same 106 tasks: the aggregate score of the harnesses ranged from 52.4 to 76.2, which is 23.8 points. A survey of June 2026 reports that on Terminal-Bench 2.0, 14 of 20 models vary by ten points or more depending on the harness, with a median range of 13.6 points for one model.

These numbers need their caveats next to them. The Harness-Bench authors call their results “configuration-level diagnostics, not causal decompositions”: the harnesses ran on default settings, part of the scoring used a model as judge, and there is one run per cell and no significance test. On the leaderboard, some gaps are inflated by broken tasks, one harness was searched on the evaluation set, and cheating has been documented. All three sources are preprints, not peer reviewed.

What it does not say

“The model is not the difference” is the strong form of the claim, and the sources that measure the effect reject it. The survey writes that the benchmark “does not support a harness-only interpretation”. Inside one harness, three OpenAI models score 35.2, 54.0 and 64.7. Harness-Bench finds that stronger backends score higher and vary less across harnesses, and its authors conclude that capability belongs to the model and the harness together.

The same paper finds the least variation across harnesses in office and business-communication tasks. Question answering over an organisation’s own knowledge sits closer to those tasks than to the long coding tasks where the gaps are largest. So we do not borrow these gaps as evidence for Inora. A measurement of our own, on our kind of work, is still to be made, and this section is where it will go.

What we claim

We say harness, and we argue the measured form: for a fixed model, the software around it can change the result a great deal, and Axiomatic builds that software for work where a wrong answer has consequences. We do not claim that the model does not matter.