When a coding agent underperforms, the reflex is to blame the model or wait for the next one. The harness, the loop, tools, context management and gates wrapped around the model, now decides more of your agent’s production behavior than the model inside it, and you engineer it rather than download it. In February 2026, LangChain lifted its coding agent 13.7 points on Terminal Bench 2.0 without touching the model. In July, a source-code audit of eleven production harnesses found four million lines with no agent framework imported anywhere. This post defines harness engineering, walks that evidence, and shows the gates we run on our own agents, including the one that caught two fabricated quotations.
The harness is everything except the model
The word needed a definition before the discipline could get a name. Birgitta Böckeler, Distinguished Engineer at Thoughtworks, published the cleanest one in April 2026: the harness is “everything in an AI agent except the model itself”. Vivek Trivedy of LangChain compressed the same idea into a formula in The Anatomy of an Agent Harness: Agent = Model + Harness, and “If you’re not the model, you’re the harness.”
Concretely, harness engineering is the design of the runtime an AI agent lives in: the execution loop, tool invocation, permissions, context and state management, error handling, and the checks that run on the agent’s output. A raw model has none of these. The agent harness is what turns a model that can describe an edit into a system that makes the edit, verifies it, and stops when it should.
Changing only the harness moved an agent 13.7 points
The cleanest controlled result so far is LangChain’s. Their February 2026 write-up on improving Deep Agents, the one linked above, states it in two sentences: “Our coding agent went from Top 30 to Top 5 on Terminal Bench 2.0. We only changed the harness.” The numbers underneath: deepagents-cli went from 52.8 to 66.5, a 13.7 point gain, across the benchmark’s 89 tasks, with the model held fixed at gpt-5.2-codex. The most common failure they fixed was not a reasoning gap. The agent would write a solution, re-read its own code, agree with itself, and stop. We saw the same self-approval pattern while putting deep agents in production.
That experiment is why Faros AI’s guide can open with the sentence this post exists to answer: “The harness, not the model, determines how well an AI coding agent performs in production.” Faros frames harness engineering as the third phase of AI engineering maturity, after prompt engineering (2022 to 2023) and context engineering (2024 to 2025). The model is a commodity you rent; the harness is the asset you own.
Eleven production harnesses, and not one agent framework
The July 2026 audit from the intro, titled Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents, read eleven production coding harnesses end to end: Claude Code, Codex CLI, Gemini CLI, Aider, OpenHands, OpenClaw and five others. Daniel Vaughan’s readable walkthrough of the 83-page study pulls out the two findings that should change how you build:
- Across roughly four million lines of Python, TypeScript and Rust, no agent runtime imports a general-purpose agentic framework. Every loop is hand-rolled.
- None retrieves code with vector embeddings. Retrieval is deterministic: ripgrep, tree-sitter, glob, and auto-discovered Markdown files.
The study catalogs 29 recurring design patterns, and the loops range from a 5,000-line linear while-loop to a 1.1 million-line async state machine, with loop sophistication predicting nothing about benchmark placement. Extension support shows where the field is heading: agent skills ship in 9 of the 11 harnesses, ahead of MCP at 8 of 11, which matters if you read our take on MCP landing inside LangChain core. Production harness authors chose predictability and debuggability over framework flexibility, in every one of the eleven codebases.
The trap: picking an agent framework first and a failure-handling strategy later. The people shipping the most-used harnesses did the opposite, and their retrieval stack is grep, not embeddings.
Is Claude Code harness engineering?
Buyers and builders type this question into Google, and the precise answer is: Claude Code is a harness; harness engineering is what you do on top of it. Out of the box it wraps a model with file tools, shell execution, a multi-step loop and permission prompts, the default rig that makes it an agent rather than a chatbot. Addy Osmani, who works on Claude Code at Anthropic, lists Claude Code, Cursor, Codex, Aider and Cline as harnesses over sometimes the same models, and gives the discipline its sharpest one-liner: “A decent model with a great harness beats a great model with a bad harness.”
The engineering starts where the defaults end. Your agents file, hooks that enforce rather than suggest, skills, subagents used as a context firewall, and the Claude Agent SDK when the harness becomes a product of its own: those extension points, not the model picker, are where teams doing serious agentic coding spend their time.
Guides steer before the act, sensors catch it after
Böckeler’s taxonomy is the most useful mental model for what to build first. Guides are feedforward controls that raise the odds of a good first attempt: rules files, templates, the shape of the repo itself. Sensors are feedback controls that observe the result and push the agent to self-correct: tests, linters, type checkers, review passes. She splits the checks themselves into computational ones, deterministic and cheap, running in milliseconds, and inferential ones such as LLM as a judge, slower, costlier and non-deterministic. Build computational sensors first; they are the only ones whose verdict you never have to double-check.
Sensors matter because self-grading fails in one direction. Prithvi Rajasekaran of Anthropic’s Labs team, writing on harness design for long-running application development, found agents reliably skew positive when grading their own work. The wider trust picture says the same: in the 2025 DORA report, 90% of nearly 5,000 respondents use AI at work, while 30% report little or no trust in the code it generates. The harness is where that gap closes. Böckeler also calls user-side harness work a specific form of context engineering, which squares with what we found when we measured context engineering in deep agents.
Every mistake becomes a permanent rule
Osmani calls the core habit the ratchet: treat every agent mistake as a permanent signal, engineer the harness so that mistake cannot recur, and let every line in your agents file trace back to a specific thing that went wrong. Constraints are added only after a real failure and removed only when a more capable model makes them redundant. The harness becomes your failure history, compiled. In that sense the rules file is the rig’s long-term agent memory, and it is also why you cannot download a good harness: it is shaped by your codebase’s failures, not the vendor’s.
Discipline beats volume here. HumanLayer’s write-up on harness engineering cites an ETH Zurich study of 138 agent context files in which LLM-generated files hurt performance while adding over 20% cost, human-written ones helped by about 4%, and agents burned 14 to 22% more reasoning tokens processing the instructions. HumanLayer keeps its own rules file under 60 lines. A rule nobody paid for with a failure is context tax.
Our own harness caught two fabricated quotations
This post was drafted, illustrated and gated inside a harness we run on our own publishing, with a person approving what ships. Two of its sensors earned their place the hard way, which is the ratchet working as described:
- A quotation gate. Every phrase a draft puts in quotation marks is matched as a literal string against the fetched bytes of the page it cites. It has caught two fabricated quotations that a structural audit had passed, each attributed to a real author beside a link to a page that never contained the words.
- A typeface gate. In one September week, two hero images shipped in a fallback serif because the renderer waited a fixed time for fonts that never arrived. The fix asks the browser which typeface actually loaded and fails the whole run otherwise, a computational sensor in Böckeler’s terms, and the defect has not shipped since.
The same loop runs our budgets: a sub-agent silent past its 30-minute ceiling is killed and replaced once, and a second silence fails the run instead of hanging it. Each of these rules traces to one dated failure, and the gains compound at the harness level too, which is how we halved Claude Code token usage by routing shell commands through a proxy, a harness change that touched no prompt and no model.
Every component is a bet that goes stale
A harness is not free, and it is not permanent. The Anthropic Labs post above is candid on both. On cost, the solo-agent run finished a project in 20 minutes for 9 dollars while the full planner, generator and evaluator harness took 6 hours and about 200 dollars, over 20 times the price, for output quality judged immediately, visibly better. On staleness: “every component in a harness encodes an assumption about what the model can’t do on its own”, and those assumptions expire. When Claude Opus 4.5 stopped rushing to finish near its context limit, the author deleted the context-reset machinery built for the model before it. Audit your harness the way the vendors audit theirs when defaults move under you: re-test each constraint against the current model, and delete the ones it has outgrown.
The takeaway
Treat the harness as the product. The model sets the ceiling, and the harness decides how close to it you operate: LangChain’s 13.7 point gain changed no weights, and the eleven most-studied harnesses in production are hand-rolled loops with deterministic retrieval and hard gates. Start with guides and computational sensors, convert every failure into a permanent rule, keep the rules file short enough that each line has a scar behind it, and re-test the whole rig when the model under it changes. If your agent disappoints, read the harness before you read the leaderboard.
Sources
- Harness engineering for coding agent users, Birgitta Böckeler
- The Anatomy of an Agent Harness, Vivek Trivedy, LangChain
- Improving Deep Agents with harness engineering, LangChain
- Harness engineering: A guide to building better AI coding agents, Faros AI
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents, Barbaste et al.
- Harness Engineering: how eleven coding agents are really built, Daniel Vaughan
- Agent Harness Engineering, Addy Osmani
- Harness design for long-running application development, Anthropic
- Skill Issue: Harness Engineering for Coding Agents, HumanLayer
- Announcing the 2025 DORA Report, Google Cloud