Harness engineering

Agentic AIArchitecture and orchestrationPublished By Simon Budziak

Harness engineering is the discipline of designing everything around an AI agent rather than the agent itself: the tools it may call, the context it receives, the checks on its output, the approval gates, and the feedback loops that let it catch its own mistakes across hours of unsupervised work.

Why did harness engineering become a named discipline in 2026?

Because teams kept learning the same lesson: agents that looked finished in a demo fell apart in production, and the fix was rarely a better model. BCG describes harness engineering as the operating system that lets agents scale safely; Thoughtworks’ Birgitta Boeckeler published a framework for it on martinfowler.com; OpenAI builds products on Codex with a harness engineering method. The industry gave the gap between a capable model and a dependable system its own name, which is usually the moment a practice becomes a budget item.

What does an agent harness actually contain?

The agent harness is the runtime: tool definitions, state, retries, and the loop the model runs inside. Harness engineering adds the design judgment around it, deciding how context is assembled for each step, which actions need a human gate, where guardrails sit, and what gets logged for the postmortem you hope never to write. A production harness is mostly checks and recovery paths, not model calls.

How is it different from prompt or context engineering?

Scope. Prompt engineering tunes a single exchange. Context engineering manages what the model knows at each step. Harness engineering owns the whole envelope across the agent lifecycle, from what the agent is permitted to do through how its work is verified. The harness, not the prompt, is where reliability is won or lost once an agent runs longer than one exchange.

What does this mean for a team buying or building agents?

Evaluate the harness, not the demo. Ask what happens on a failed tool call, who approves consequential actions, how a run is traced afterward, and how confidence gates decide when the agent must stop and ask. Two agents on the same model can differ by an order of magnitude in reliability, and the difference is engineering you can inspect.

This entry was drafted with AI assistance.

Frequently asked questions

What is harness engineering in AI?

The practice of building the execution environment around an AI agent: tool access, context assembly, state, guardrails, approvals, retries, and monitoring. The model supplies the reasoning; the harness decides what that reasoning can touch and how failures get caught.

Is harness engineering the same as prompt engineering?

No. Prompt engineering shapes one exchange with a model. Harness engineering shapes whether an agent stays reliable across hundreds of decisions, which depends far more on tools, checks, and recovery paths than on the wording of any single prompt.

Is Claude Code an example of harness engineering?

Yes, coding agents are the clearest examples: the model is paired with a harness of file access, shell execution, test runners, and permission prompts, and most of the difference between such products lives in the harness, not the underlying model.

Summarize this page with

See this working in a system we built