Harness Engineering
AI models can already write a million lines of code. The real challenge is no longer making them write better, buthow to harness them to work stably, reliably, and without losing control. This methodology of building constraints, feedback, and control systems around AI agents is the new paradigm that rapidly swept through the engineering world in early 2026 —Harness Engineering。
1. What is Harness Engineering?
Harness Engineeringis the systematic engineering practice of designing and buildingconstraint mechanisms, feedback loops, workflow control, and continuous improvement cyclesaround AI agents.
It does not optimize the model itself, but the environment in which the model operates. The core philosophy boils down to one phrase:Human Steer, Agent Execute。
The word "harness" comes from horse tack — reins, saddle, bit — a complete set of equipment used to guide powerful but unpredictable animals.Harness Engineering is not about weakening AI's capabilities, but about forging a golden harness for it so it can run fast and steady.
This concept was first proposed by HashiCorp co-founder Mitchell Hashimoto on February 5, 2026. Six days later, OpenAI formally adopted the term in its one-million-line code experiment report, and Martin Fowler subsequently wrote an in-depth analysis. Within a month, it became a buzzword in the developer community.
harness engineering is the idea that anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent will not make that mistake again in the future.
—— Mitchell Hashimoto
The subtext of this statement is:Every failure of an Agent is a signal that the environment design is incomplete.The correct response is not to switch to a more powerful model, but to redesign the environment it operates in.
2. Why Do We Need Harness Engineering? Real Data Speaks
The LangChain case is especially compelling: not a single parameter of the underlying model was changed;merely by optimizing the external steering environment(documentation structure, verification loops, tracking system), the coding Agent's score on Terminal Bench 2.0 soared from 52.8% to 66.5%, and its global ranking jumped from 30th to 5th.
Five independent teams arrived at the same conclusion:The bottleneck is not model intelligence, but infrastructure.
3. Three Paradigm Shifts in AI Engineering
To understand why Steering Engineering matters, we first need to see how we got here step by step.
| Paradigm | Core problem | Optimization target | Interaction mode |
|---|---|---|---|
| Prompt Engineering | How to express things clearly | Prompt wording, format, examples | Question-and-answer |
| Context Engineering | How to feed information to AI | Documents, code snippets, conversation history | Information injection → generation |
| Steering Engineering | How to make Agents work reliably | Constraints, feedback loops, control systems | Humans steer, Agents execute |
A memorable analogy:
- Prompt Engineering —— The skill of shouting commands to a horse
- Context Engineering —— The map shown to the horse
- Harness Engineering —— Building a highway for the horse, complete with guardrails, speed limit signs, and gas stations
4. Common Agent Failure Modes
Through long-running agent operations, Anthropic engineers summarized three typical failure modes—precisely the core pain points that harness engineering must solve:
Failure Mode 1: One-shotting
Agents tend to implement all functionality in a single session. The result is a depleted context window, a pile of undocumented half-finished code, and the next session must spend a lot of time guessing what happened before.
Failure Mode 2: Prematurely declaring victory
Late in a project, once some features are done, the agent looks around and, seeing progress, immediately declares the task complete—even though many features remain unimplemented.
Failure Mode 3: Prematurely marking features as complete
Without explicit prompting, the agent marks code as complete right after writing it, without end-to-end testing. Passing unit tests or curl commands does not mean the feature actually works.
Additionally, agents have another dangerous trait:They are extremely good at pattern copying. It faithfully replicates and amplifies whatever patterns exist in the codebase—including bad patterns and architecture drift. This means an unconstrained agent accumulates technical debt at an alarming rate.
5. The Four Guardrails of Harness Engineering
Drawing on the practices of OpenAI, Anthropic, LangChain, and Martin Fowler, Harness can be summarized into four core components—four "guardrails":
Guardrail 1: Context Engineering — The New Employee Handbook
Just like giving a new employee a detailed work manual,AGENTS.mdit is the first guide an AI agent sees when entering a codebase. But this is not a static 1,000-page manual—context is a scarce resource, and too much guidance will instead crowd out space for tasks, code, and related documentation, turning it into a graveyard of stale rules.
A better approach is: provide a stable, compact entry point, then teach the Agent to retrieve and pull in more context as needed based on the current task. In Mitchell Hashimoto's Ghostty project, every line in AGENTS.md corresponds to a historical Agent failure case—the documentation is a living feedback loop, not a static artifact.
Guardrail 2: Architecture Constraints — The Reins
The OpenAI team established a strict hierarchical dependency model:
Types → Config → Repo → Service → Runtime → UI
Lower layers cannot have reverse dependencies on upper layers. All architecture rules are encoded asCustom Linter Rules, violating it will make CI block the merge—whether the code is written by humans or AI.
There's a key detail: the Linter's error messages themselves are also context engineering. It doesn't just say you violated rule X, but explains why the rule exists and what the correct approach is, so that when the Agent reads the error, it can understand and correct itself without human intervention.
Guardrail 3: Feedback Loop — Agents Reviewing Agents
In traditional development, human engineers are responsible for code review. In agentic engineering, this work becomes agent-to-agent: Codex reviews its own changes locally, requests additional review, and loops until it passes.
Hooks in the feedback loop can run predefined test suites and, on failure, loop back to the model with error information, or prompt the model to independently evaluate its code. If AI-written test cases pass code that contains bugs, the Harness will deem the test invalid, forcing it to rethink the test boundaries.
Guardrail 4: Entropy Management — Garbage Collection
Over time, software systems gradually become chaotic (entropy increases), and technical debt accumulates. OpenAI adopts a strategy of continuous small repayments instead of waiting until problems become serious to handle them centrally—they vividly refer to this approach asGarbage recycling, and believes that technical debt is like a high-interest loan.
Specific measures: regularly run background Codex tasks to scan for deviations, update quality ratings, and initiate targeted refactoring PRs. In addition, there is also a dedicated...Doc-gardening Agent(Document Gardener Agent) automatically scans in the background for inconsistencies between documentation and code, and when it finds outdated content, it automatically submits a PR to fix it — agents maintaining documentation for agents.
6. Six Industry-Wide Consensus Points
Based on multiple independent information sources including OpenAI, Anthropic, LangChain, Stripe, and HashiCorp, the industry has formed a clear consensus on the following six aspects:
| # | Consensus | Core viewpoint |
|---|---|---|
| 1 | The bottleneck is in infrastructure, not in model intelligence. | Five independent teams reached the same conclusion. Just changing the Harness tool format can jump the model's score from 6.7% to 68.3% |
| 2 | Documentation must be a living feedback loop | Static documentation is a graveyard; only dynamic documentation has value. Let background Agents periodically clean up outdated docs and submit PRs |
| 3 | Separation of thinking and execution | Complex tasks cannot be completed in a single context window; an Orchestrator + Worker layered architecture is needed, with state persisted to external storage |
| 4 | More context is not always better | Context is a scarce resource. Huge instruction files crowd out task space; retrieve on demand and inject dynamically |
| 5 | Constraints must be automated | Manual review is the bottleneck. Guardrails should be encoded as Linter, CI, and type systems, letting machines enforce them instead of humans |
| 6 | The engineer's role is shifting | From code writer to environment architect. The biggest engineering challenge is designing a control system that makes Agents work reliably |
7. The Relationship Between Harness and Traditional Frameworks
Harness is not a replacement for SDKs, scaffolds, or Agent frameworks, but rathera layer on top of them:
Traditional frameworks solvehow to build AI agents, while the Harness Layer solves a completely different problem:how to run agents reliably。
Models are gradually absorbing about 80% of framework functionality (agent definition, message routing, task lifecycle...), but the remaining 20% —persistence, deterministic replay, cost control, observability, error recovery— is exactly the value of the Harness Layer.
Summary
Harness Engineering is not an experiment by one company, but a paradigm shift the entire industry is undergoing.
Birgitta Böckeler's summary is the most insightful:
To achieve higher AI autonomy, the runtime must be subject to stricter constraints. Increasing trust requires not more freedom, but more restrictions.
Like guardrails on a highway — it is precisely because of the guardrails that you dare to push it to 120 km/h.
| Core components | Problems solved | Representative practices |
|---|---|---|
| Context engineering Context | Agent doesn't know what to look at or how to find it. | AGENTS.md: living document, on-demand retrieval. |
| Architectural constraints Constraints | Agent copies and amplifies bad patterns. | Layered dependencies, custom Linter, CI-enforced blocking |
| Feedback loop Feedback | Agent does not know that it did something wrong. | Agent-to-Agent Review, Automated Test Suite |
| Entropy management Entropy | Technical debt and documentation rot. | Doc-gardening Agent, continuous garbage collection |
The future of software development may no longer be about how fast and well we can write code, but rather aboutHow smart and robust systems can we design to harness the immense power of AI agents?。
The value of engineers is shifting from executors to enablers and systems thinkers — from building products to building factories that can build products.