Agent Native
Harness Engineering
Agent Native·A book for AI engineers

Harness Engineering

Two teams can call the same model and get very different agents. The difference is the harness: the loop that drives the model, the context it assembles on every turn, the tools it exposes, the checks that push back on bad work, the sandbox that bounds what an action can touch, and the state that lets work outlast one context window. This book takes each part in turn, from the primary sources and measured results available in October 2026, with code you can run and figures you can step through, and then turns to the repository itself as the harness your coding agents work in.

13 chapters + appendix·Interactive figures and runnable code·Current to October 2026·Agent Native, 2026

Chapter 1 · The case

The Harness Is What You Ship

One model wired to three harnesses with their Terminal-Bench 2.0 scores

What the measured spreads between harnesses show, where they shrink to a few points, what the word covers, and the habit that turns each agent failure into a permanent fix.

On the Terminal-Bench 2.0 leaderboard snapshot of August 28, 2026, Claude Opus 4.6 appears in nine harnesses. With the same weights and the same 89 tasks, it scores 57.98% in Claude Code, at rank 50, and 76.40% in Meta-Harness, at rank 11. That is an 18.4-point spread with the model held constant.

Only two of those nine rows carry the board's verified mark, Terminus 2 at 62.9% and Claude Code at 58.0%, and between them the spread is 4.9 points. The seven higher entries are unverified third-party submissions. The entry that once topped them all, ForgeCode at 81.8%, fell to about 71.7% in Meerkat's audit at UPenn once the runs in which its agent curled solutions into its AGENTS.md were swapped for clean ones, and it is no longer on the board.

The first number shows how much room the harness has, and the second how little of the most quoted evidence survives a check. This chapter weighs the evidence and the counter-evidence, defines the term, traces its origin, maps the subsystems the book builds, and ends with two habits: fixing each agent failure where it cannot recur, and reading every harness claim by its model, harness and benchmark version.

Who this book is for and what you will build

This book is for AI engineers who build agents and for engineers who run coding agents such as Claude Code, Codex, OpenCode or Deep Agents on production codebases. The first group writes a harness. The second configures one that someone else wrote, through instruction files, hooks, permission rules, skills and the repository the agent works in. Section 4 calls these the builder harness and the user harness, and the book shows every mechanism both ways: as code you own, and as the Claude Code or Codex configuration that does the same job.

The code you will build is hx, a small Python package that grows by one module per chapter. Chapter 2 writes the core: a run() loop over the Anthropic Messages API, the Tool and State types, an event log that can rebuild the exact request sent on any turn, and four middleware hooks (before_request, before_tool, after_tool, before_stop), the last three mirroring Claude Code's PreToolUse, PostToolUse and Stop events. Each later chapter plugs one module into those hooks, and Chapter 13 assembles them into a background agent that takes an issue to a reviewed pull request, with its unit economics. The code targets Python 3.11 or later and the anthropic SDK, with claude-sonnet-5-5 as the default model.

The chapters follow the subsystem map in section 6:

  • Chapters 1 and 2: the case, and the loop one turn at a time.
  • Chapters 3 to 5: what the model sees and does (context; caching, compaction and memory; tools).
  • Chapters 6 and 7: what checks and constrains it (feedback; permissions, sandboxes and the network).
  • Chapters 8 and 9: work across agents and across hours (orchestration; long-running work).
  • Chapters 10 and 11: the environment and the team (the repository; review and merge).
  • Chapters 12 and 13: measuring and evolving a harness, then one built end to end, plus an appendix of sources.

From this chapter you take two programs and a habit. first_turn_audit.py measures what a harness spends before the model reads your task, ledger.py keeps the failure ledger later chapters build on, and the habit in section 8 is Mitchell Hashimoto's: when the agent makes a mistake, change the harness so it cannot make that mistake again.

Same model, different scores

A leaderboard that accepts many harnesses shows the effect directly, because the same model appears in several rows. The Terminal-Bench 2.0 paper (Merrill, Shaw, Carlini and 82 coauthors) adds a controlled version: its Table 2 runs each model in several harnesses on the same 89 tasks, with 95% intervals of two to three points. GPT-5 scores 49.6% in Codex CLI and 33.9% in Mini-SWE-Agent; Gemini 2.5 Pro scores 32.6% in Terminus 2 and 15.7% in OpenHands. Claude Sonnet 4.5 stays between 40.1% and 42.8% in all four harnesses it ran in.

Pick a model to see its rows. Opus 4.6 and Haiku 4.5 come from the board snapshot of August 28, 2026; GPT-5, Gemini 2.5 Pro and Sonnet 4.5 come from the paper's table, where the authors ran every pair. Turn on VERIFIED ONLY and the board's unverified rows go to hatch: Opus 4.6's spread falls from 18.4 points to 4.9, while Haiku 4.5's falls only from 21.6 to 15.9. Whiskers mark the board's 95% interval where it records one.

Leaderboard rows also bundle in configuration, infrastructure, trial count and date. Stronger evidence holds the model fixed and changes only the harness. In “Improving Deep Agents with harness engineering” LangChain moved deepagents-cli from 52.8% to 66.5% on Terminal-Bench 2.0 with the model fixed at gpt-5.2-codex, from just outside the top 30 to the top 5 (“We only changed the harness.”). The reasoning budget was one of the harness settings. Running at xhigh throughout scored 53.9% because runs hit timeouts, high throughout scored 63.6%, and the final 66.5% came from an xhigh, high, xhigh “reasoning sandwich”.

Sydney Lewis's “Same Model, Different Harness” changes less: one harness, run with the full transcript and with old tool results mechanically shortened plus stall handling. On 169 SWE-bench Verified tasks at a 20,480-token window, the mean fail-to-pass fraction rose from 28% to 49% and complete solutions from 43 to 72, and the abstract reports the effect holding for three more models without retuning. Its conclusion is this book's premise: “coding-agent evaluations should treat the model and harness together as the tested solver.”

Agentic Harness Engineering (AHE) hands the tuning to an agent. Ten evaluate, analyze and improve iterations over about 32 hours, with GPT-5.4 held at high effort, lifted a bash-only seed from 69.7% to 77.0% on Terminal-Bench 2, past the hand-built OpenCode (47.2%), Terminus-2 (62.9%) and Codex (71.9%) harnesses on the same model. Swapped into the seed one at a time, the evolved memory added 5.6 points, the tools 3.3 and the middleware 2.2, while the evolved system prompt alone cost 2.3. Those three gains sum to 11.1 points, yet the full harness gains 7.3.

Can Bölük's “The harness problem” changed a single tool. Replacing the patch edit format with “hashline”, which tags every line the model reads so that edits can point at the tags, beat patch for 14 of 16 models by 15 points on average, and Grok Code Fast 1 went from 6.7% to 68.3%. The experiment cost about $300 of benchmarking. Chapter 5 covers edit formats.

The production evidence is a company self-report. Cline moved its VS Code extension, with more than 11 million users, from a roughly 76,000-line monolith to its SDK harness. It ran the old and new builds near 50/50 behind a feature flag for over a week and reports that tasks ending in task.mistake_limit_reached (three consecutive mistakes) fell from 6.34% to 0.62%. Its explanation is an assumption that expired: the 2024 harness parsed XML tool calls out of the text stream, and “Models in 2026 are RL-trained to call tools natively.”

Independent measurement points the same way. Epoch AI found that switching scaffold on SWE-bench Verified makes a difference of up to 11% for GPT-5 and up to 15% for Kimi K2 Thinking, and wrote that “The choice of scaffold has the single biggest impact on the overall performance.”

StudyHeld fixedBenchmarkWhat changedResultGrade
LangChain, Feb 2026gpt-5.2-codexTerminal-Bench 2.0, 89 tasksprompt, tools, middleware, reasoning budget52.8% to 66.5%primary; a vendor tuning its own agent
Bölük, Feb 2026each of 16 models180 edit tasks on React files, 3 runsthe edit tool onlyGrok Code Fast 1: 6.7% to 68.3%primary
AHE, Apr 2026GPT-5.4, high effortTerminal-Bench 2, 89 tasksan agent evolves the harness, 10 iterations69.7% to 77.0%primary
Lewis, Aug 2026one model, three more checked169 SWE-bench Verified tasks, 20,480-token windowold tool results trimmed, stall handlingfail-to-pass 28% to 49%primary (abstract)
Fan et al., Sep 2026Nemotron-3 550BSWE-bench Verified, 32k budgetsummarization against no context management6.40% to 58.40%primary
Cline, Sep 2026five production modelsproduction A/B, 11M+ userslegacy harness to SDK harnessmistake-limit stops 6.34% to 0.62%company self-report

The studies measure different things on different scales, so the next figure gives each one its own axis, with the starting harness hatched and the changed one in white.

Switch between the studies to compare them. The Terminal-Bench views share a 0 to 100% axis, Cline's is a failure rate where shorter bars are better, and the CONTEXT view, which the next section picks up, plots how many points context management adds at each token budget.

How large the effect is, and when it shrinks

The Terminal-Bench authors read their own table the other way. Moving Gemini 2.5 Pro from OpenHands to Terminus 2 is worth 16.9 points (15.7% to 32.6%), while moving Codex CLI from GPT-5-Nano to GPT-5.2 is worth 51.4 (11.5% to 62.9%), and they conclude that “model selection is usually more important than agent scaffold when optimizing for performance.” Both readings hold. The harness effect is measurable, and the model effect is usually larger.

Several careful studies find the harness effect on success small, or absent, for strong models:

  • METR measured time horizons inside the product harnesses and found that “neither Claude Code nor Codex outperform the default scaffolds METR uses.” Claude Code beat METR's ReAct scaffold in 50.7% of bootstrap samples for Opus 4.5, and Codex beat Triframe in 14.5% for GPT-5; neither difference was significant.
  • Scale and Reflection's SWE-Bench Pro V2 harness study ran GLM-5.3, Kimi K3 and Inkling through mini-swe-agent, Pi and OpenCode on 642 tasks under a locked budget and found them on par for the strongest models, except Inkling on Pi. Reflection's advice: “Pick the harness for the bill, not the score.”
  • Arena.ai's HarnessTax study, which this book's research grades as secondary, reports that across 21 model and harness pairs the harness moved success by no more than about 2% on SWE-bench Lite and about 5% on Terminal-Bench 2.0, while Claude Code cost about 2.0 times as much as Pi on SWE-bench Lite.
  • Qwen Code's regression table holds Qwen 3.7 Max fixed across seven harness versions, and the average on SWE-bench Verified stays between 76.40% and 77.80%.

The effect concentrates where the model has less room: weaker or cheaper models, tight context budgets and long terminal tasks. In the Terminal-Bench 2.0 paper's table the spread across harnesses is 15.7 to 16.9 points for GPT-5, Haiku 4.5 and Gemini 2.5 Pro, and 2.7 to 3.2 points for Sonnet 4.5 and Opus 4.1. The clearest dose-response comes from Fan and colleagues' empirical study of harness design, which ran a fixed ReAct loop over 176 matched settings and compared them with exact McNemar tests. Context management was worth 35.7 points on SWE-bench Verified at a 32k token budget, 15.9 at 64k, 5.5 at 96k and 2.7 at 128k, and on Terminal-Bench 2.1 it fell from 9.5 to 2.8; without it, 78.7% of SWE-bench runs overflowed at 32k. The CONTEXT view of the figure above draws the decline. In the same vein, “Beyond the Model” reports that complex harnesses give diminishing gains on SWE-style repair as models improve, while still helping stronger models on open-ended repository tasks.

The benchmark moves under the harness too. Terminal-Bench 2.1 fixed 28 of the 89 tasks, nine of them broken by external dependency drift and eight by resource mismatches. Opus 4.6 in Claude Code rose from 58.0% to 70.1%, the largest gain of any pair, while Opus 4.6 in Terminus 2 moved from 62.9% to 63.8%, so broken tasks explain part of why, in LangChain's words, “The Claude Code harness ranks last among Opus 4.6 submissions” on 2.0. Chapter 12 covers benchmark versions and noise.

Before you quote a leaderboard spread
Check the verified flag and the snapshot date, then the integrity notes. Terminal-Bench's integrity policy records three cases: OB-1 stored encrypted solutions in its agent binary, Pilot uploaded the tests folder, and ForgeCode's agent curled solutions into its AGENTS.md. Meerkat's audit found the top three Terminal-Bench 2 submissions guilty of cheating and a Meta-Harness Opus 4.6 trace that printed “PASS” to fool a verifier. Meerkat argues that harness-level cheating is often “meta” reward hacking by the coding agents that developers use to build their harnesses.

Where success barely moves, cost still does. Systima compared the first request two harnesses send on claude-sonnet-4-5 in July 2026. Claude Code 2.1.207 sent about 32,800 tokens before the model saw the task, from a 27,344-character system prompt and 27 tools whose schemas ran to 99,778 characters; OpenCode 1.17.18 sent about 6,900, with 10 tools. Claude Code wrote up to 54 times more cache tokens, and a two-subagent fan-out turned a 121K-token task into 513K tokens.

You can take the same measurement of any harness that speaks the Anthropic Messages API. The first script stands in for the API on localhost and saves the request bodies a harness sends; the second counts what each layer of the largest request costs.

capture_requests.py
python
#!/usr/bin/env python3
"""Record the Messages API requests a harness sends, then refuse them.

1. python capture_requests.py            (listens on 127.0.0.1:8787)
2. In your repository, in a second terminal, start the harness against it:
       ANTHROPIC_BASE_URL=http://127.0.0.1:8787 ENABLE_TOOL_SEARCH=true claude
   and send one short message. The harness reports an API error; that is expected.
3. python first_turn_audit.py captured/

Bodies are saved to captured/NNN.json; headers, which carry your key, never are.
A proxy URL turns Claude Code's MCP tool search off by default, and
ENABLE_TOOL_SEARCH=true keeps MCP tools deferred, as in a direct session.
"""
import json
from http.server import BaseHTTPRequestHandler, HTTPServer
from pathlib import Path

OUT = Path("captured")
REFUSAL = json.dumps({"type": "error", "error": {
    "type": "invalid_request_error", "message": "captured by capture_requests.py"}}).encode()


class Capture(BaseHTTPRequestHandler):
    def read_body(self) -> bytes:
        if self.headers.get("transfer-encoding", "").lower() == "chunked":
            body = b""
            while size := int(self.rfile.readline().split(b";")[0].strip() or b"0", 16):
                body += self.rfile.read(size)
                self.rfile.readline()  # CRLF after each chunk
            return body
        return self.rfile.read(int(self.headers.get("content-length", 0)))

    def do_POST(self) -> None:
        body = self.read_body()
        if self.path.split("?")[0].endswith("/v1/messages"):
            OUT.mkdir(exist_ok=True)
            path = OUT / f"{len(list(OUT.glob('*.json'))):03d}.json"
            path.write_bytes(body)
            print(f"saved {path} ({len(body):,} bytes)")
        self.send_response(400)
        self.send_header("content-type", "application/json")
        self.send_header("content-length", str(len(REFUSAL)))
        self.end_headers()
        self.wfile.write(REFUSAL)

    def log_message(self, *args) -> None:  # silence the default access log
        pass


HTTPServer(("127.0.0.1", 8787), Capture).serve_forever()
first_turn_audit.py
python
#!/usr/bin/env python3
"""Count what a harness spends on its first request, layer by layer.

    python first_turn_audit.py captured/                  # the largest captured request
    python first_turn_audit.py request.json --model claude-sonnet-4-5 --window 200000

Each layer is counted as the tokens it adds to a one-line request, with the
Messages API token counter (needs ANTHROPIC_API_KEY). Any request body in
Messages API format works: a capture, or what hx's rebuild_request() returns.
"""
import argparse
import json
from pathlib import Path

import anthropic

client = anthropic.Anthropic()
PING = [{"role": "user", "content": "hi"}]


def count(model: str, **parts) -> int:
    parts = {k: v for k, v in parts.items() if v}  # drop empty layers
    parts.setdefault("messages", PING)
    return client.messages.count_tokens(model=model, **parts).input_tokens


def main() -> None:
    ap = argparse.ArgumentParser()
    ap.add_argument("path", type=Path)
    ap.add_argument("--model", help="defaults to the model named in the request")
    ap.add_argument("--window", type=int, default=200_000)
    args = ap.parse_args()

    files = sorted(args.path.glob("*.json")) if args.path.is_dir() else [args.path]
    req = max((json.loads(f.read_text()) for f in files), key=lambda r: len(r.get("tools", [])))
    model = args.model or req["model"]

    system = req.get("system") or None
    if isinstance(system, list):  # keep the text, drop cache_control and other keys
        system = [{"type": "text", "text": b["text"]} for b in system if b.get("type") == "text"]
    # The counter rejects server tools (web search, tool search); deferred tools are not in the prompt.
    tools = [{k: t[k] for k in ("name", "description", "input_schema") if k in t}
             for t in req.get("tools", []) if "input_schema" in t and not t.get("defer_loading")]
    first = next(m for m in req["messages"] if m["role"] == "user")
    content = first["content"] if isinstance(first["content"], list) else [{"type": "text", "text": first["content"]}]
    blocks = [b["text"] for b in content if b.get("type") == "text" and b["text"].strip()]

    base = count(model)
    rows = [("system prompt", count(model, system=system) - base)] if system else []
    single = [(t["name"], count(model, tools=[t]) - base) for t in tools]
    overhead = 0
    if len(tools) > 1:  # every single-tool count repeats any fixed tool-use preamble
        together = count(model, tools=tools) - base
        overhead = max(0, (sum(n for _, n in single) - together) // (len(tools) - 1))
        rows.append(("tool-use overhead", overhead))
    rows += [(f"tool {name}", n - overhead) for name, n in single]
    for i, text in enumerate(blocks):
        head = text.strip().splitlines()[0][:44]
        rows.append((f"user block {i}: {head}", count(model, messages=[{"role": "user", "content": text}]) - base))

    user = [{"role": "user", "content": [{"type": "text", "text": t} for t in blocks]}] if blocks else PING
    total = count(model, system=system, tools=tools, messages=user)
    width = max((len(name) for name, _ in rows), default=10) + 2
    print(f"{'layer':<{width}}{'tokens':>9}{'share':>8}")
    for name, n in sorted(rows, key=lambda r: -r[1]):
        print(f"{name:<{width}}{n:>9,}{n / total:>8.1%}")
    print(f"{'first request':<{width}}{total:>9,}  {total / args.window:.1%} of a {args.window:,}-token window")
    if skipped := len(req.get("tools", [])) - len(tools):
        print(f"{skipped} server or deferred tools not counted")


if __name__ == "__main__":
    main()

Each row is the tokens one layer adds to a one-line request, so the rows sum to roughly the total, and instruction files such as CLAUDE.md and AGENTS.md show up in whichever row the harness puts them. To compare with Systima's figures, pass --model claude-sonnet-4-5; their measurement put Claude Code's tool schemas at more than three times the size of its system prompt. Chapters 3 to 5 cover what to do with a heavy row, and for shrinking what a long session carries, see our field guide to context compression.

What the word covers: builder harness and user harness

“Harness” replaced “scaffold” in about a year. Anthropic's January 2025 SWE-bench post defined an agent as “the combination of an AI model and the software scaffolding around it.” Its January 2026 post on evals writes “agent harness (or scaffold)” and keeps it apart from the evaluation harness that runs the tests, and the Claude Code glossary now says: “Claude Code is the harness; Claude is the model inside it.” The definitions in circulation agree on the boundary and differ on what they emphasize.

SourceIn their wordsSide it describes
LangChain, Vivek Trivedy, Mar 2026“Agent = Model + Harness.” “If you're not the model, you're the harness.” “A harness is every piece of code, configuration, and execution logic that isn't the model itself.”both, written from the builder's side
Anthropic, Lance Martin, Apr 2026“the software scaffolding around a model: the loop, tools, context management, and guardrails”; design is “deciding what belongs in that scaffolding and, as models improve, what you can take out”builder
Claude Code docs, glossary, Oct 2026“The tools, context management, and execution environment that turn a language model into a capable coding agent.”builder
Birgitta Böckeler, martinfowler.com, Apr 2026“everything in an AI agent except the model itself”; three bounded contexts: the model, the coding agent's builder harness, and the user harness as an outer ringnames both
Dex Horthy, HumanLayer, Nov 2025“applying context engineering principles to how you use an existing agent”; the commands, hooks, skills, agents and MCPs a user plugs inuser
Barbaste et al., arXiv 2609.00006, Jul 2026“the design and evolution of that runtime”; a harness is not quite a scaffold, not an agentic framework and not an evaluation harness (“Same word, opposite direction of wrapping.”)builder, as the shipped runtime

The split matters because two people can say “harness engineering” and mean different jobs. OpenAI's “Harness engineering” post is about the user and repository side, the docs, linters and observability that let Codex work in a codebase, while Anthropic and LangChain mostly describe the builder side: the loop, tools and compaction. Most sources never name the split, and readers talk past each other. Böckeler's article names it and files the user side under a larger idea: “Engineering a user harness for a coding agent is a specific form of context engineering.”

Which of the two contains the other is unsettled. Horthy, HumanLayer and Böckeler treat harness engineering as a form of context engineering, and LangChain calls today's harnesses “delivery mechanisms for good context engineering.” Databricks writes the reverse, “Prompt and context engineering both live inside harness engineering,” and Andrej Karpathy, in his June 2025 post endorsing the term context engineering, already called it “one small piece of an emerging thick layer of non-trivial software.”

This book uses three terms. The harness is everything in a running agent that is not the weights. The builder harness is the code of the loop, its tools and its middleware: hx in this book, or Claude Code's own source. The user harness is what you add around a harness you did not write: CLAUDE.md and AGENTS.md files, hooks, permission rules, skills, MCP servers, and the repository's docs, lints and tests. “Evaluation harness” always means the thing that runs a benchmark, and “scaffold” appears only in quotations. Section 8 shows one rule in both forms.

Where the term came from

The earliest documented use of the exact phrase is a post on X by Dex Horthy of HumanLayer on November 4, 2025. He introduced a concept “which I'll call ‘harness engineering’”, applying context engineering principles to how you use an existing agent, and for Claude Code he listed “the commands, hooks, skills, agents, mcps, etc that a consumer of the tool plugs into the existing harness.” Hacker News has no story or comment with the phrase in any month from January to November 2025, and the first comment appears on December 19, 2025.

From scaffold to harness engineering

  1. Scaffolding around the model
    2025-01-06
    Anthropic's SWE-bench post defines an agent as “the combination of an AI model and the software scaffolding around it” and advises keeping the scaffolding minimal.
  2. Context engineering
    2025-06-19 to 2025-06-25
    Tobi Lütke prefers “context engineering” to prompt engineering; Andrej Karpathy adds that it is one piece of a “thick layer” of software around model calls.
  3. “Agent Harness” and the Claude Agent SDK
    2025-09-23 to 2025-09-29
    Vivek Trivedy defines an agent harness and coins “Harness as a Service”, without the phrase “harness engineering”. Anthropic renames the Claude Code SDK the Claude Agent SDK, “the agent harness that powers Claude Code”.
  4. Horthy names harness engineering
    2025-11-04
    The earliest documented use of the exact phrase, in a post on X.
  5. “The Codex harness”
    2026-01-23 and 2026-02-04
    OpenAI's Michael Bolin and Celia Chen describe “the Codex harness”, the agent loop and logic under every Codex surface.
  6. Hashimoto: Engineer the Harness
    2026-02-05
    Step 5 of “My AI Adoption Journey”: engineer each mistake away so the agent never makes it again. The post reaches 984 points on Hacker News.
  7. OpenAI's “Harness engineering”
    2026-02-11
    Ryan Lopopolo's post on building a product with no hand-written code. It never defines the term and never claims to have coined it.
  8. Bölük, LangChain, Böckeler
    2026-02-12 to 2026-04-02
    “The harness problem” (832 points on Hacker News); LangChain's “Improving Deep Agents with harness engineering”, Trivedy's first located use of the phrase; Böckeler's memo, then her article on the builder and user harness.
  9. A track, a job title, a product category
    2026-04 to 2026-09
    “Agent Harness Engineer” job postings; a Harness Engineering track at the AI Engineer World's Fair, where 18 of 561 session titles say “harness” and 2 say “context engineering”; harness products from Vercel, Microsoft and OpenAI; a Wikipedia article that calls the attribution contested.

Published attributions disagree, and none of the prominent ones cites Horthy. HumanLayer, in March 2026, and Addy Osmani, in April, credit Vivek Trivedy, whose first located use of the exact phrase is LangChain's post of February 17, 2026; his 2025 post defines an agent harness but never says “harness engineering.” The arXiv source-code study (2609.00006) says Hashimoto used it first and Trivedy then defined it, citing a secondary source, and elsewhere calls it “a phrase coined in a vendor blog post.” Wikipedia calls the attribution contested between those two. Horthy co-founded HumanLayer, the company that credits Trivedy.

OpenAI's post, which InfoQ covered as “a new internal engineering methodology called Harness engineering,” never claims the coinage, and Hashimoto's use came six days earlier. A quotation about “Harness” that InfoQ attributed to Ryan Lopopolo does not appear in the post, so this book does not repeat it. We write “earliest documented use” and cite Horthy, because earlier uses may exist.

The words for the next layer up have already appeared. On June 7, 2026 Peter Steinberger wrote that “You should be designing loops that prompt your agents,” and Osmani's “Loop Engineering” placed itself “one floor above the harness.” The appendix tracks how the term spread.

The subsystem map

Production harnesses converge on one shape. A source-code study of 11 production harnesses (Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode and OpenClaw), about 4 million lines in all, found that none imports a general agent framework or retrieves code with embeddings. SKILL.md appears in nine of them and MCP in eight, and the authors derive a 90-line minimum harness; our teardown of that study walks through its patterns. VILA Lab's reading of the Claude Code v2.1.88 source puts “AI decision logic” at 1.6% of the codebase, which leaves the rest to the harness.

The map below is the one this book follows: the model in the middle, the loop around it, the parts that feed, act for and check that loop, and the repository, review and measurement around all of them.

Hover over a part, or tab through them, to read its job, what breaks without it, and the chapter that builds it. Chapter 13 assembles every part into one agent.

The parts meet inside every request. OpenAI's “Unrolling the Codex agent loop” gives the order in which Codex assembles one: the model's instructions, the tools, a permissions message, the AGENTS.md files from the project root down to the working directory (up to 32 KiB), an environment-context message, and only then the user's message. Every turn resends that prefix, which makes the loop quadratic in JSON sent, and only exact-prefix cache hits make sampling “linear rather than quadratic.” Changing the tools, model, sandbox, approval mode or working directory mid-conversation breaks the cache, so context, permissions, tools and caching share one constraint. Chapter 4 works through the arithmetic.

The scorecard turns the map into an audit of your own setup. Each item is the smallest working version of one subsystem, and its note names the chapter that builds it.

Audit your harness against the map

0% (Foundational)

Every component encodes an assumption

Prithvi Rajasekaran of Anthropic Labs stated the rule in a March 2026 post on harness design for long-running application development: “every component in a harness encodes an assumption about what the model can't do on its own,” and those assumptions “can quickly go stale as models improve.” The space of interesting harness combinations, he adds, “doesn't shrink as models improve. Instead, it moves.”

Anthropic's own harnesses show it. Sonnet 4.5's “context anxiety” made compaction alone insufficient, so the long-running harness reset the context and passed handoff files between sessions. With Opus 4.5, in the words of the Managed Agents post, “the behavior was gone. The resets had become dead weight,” and the harness moved to the Agent SDK's automatic compaction. For Opus 4.6 the sprint construct went as well and the evaluator moved to a single pass at the end, but the planner stayed, because without it the generator under-scoped. Assumptions can also flip back: Amp replaced compaction with handoffs in October 2025 and restored it in May 2026, because “Today's leading frontier models are great at handling compaction.”

Claude Code is the best-documented case. As reported in the Pragmatic Engineer's interview with Boris Cherny in September 2025, “with the 4.0 models, we deleted around half the system prompt.” In July 2026 Anthropic wrote that “We removed over 80% of Claude Code's system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations.” The post, by Thariq Shihipar, says “we were overconstraining Claude Code, both through our system prompt and in our CLAUDE.md files and skills,” and trades rules for judgment, examples for interfaces and upfront instructions for progressive disclosure. One old line, “In code: default to writing no comments. Never write multi-paragraph docstrings...”, became “Write code that reads like the surrounding code: match its comment density, naming, and idiom.”

The model-facing tools shrank too. Version 2.1.117 took Glob and Grep out of native macOS and Linux builds in favor of search through the Bash tool, and version 2.1.233 stopped offering the to-do and task tools to Opus 4.8, Sonnet 5, Fable 5, Mythos 5 and newer models. The product around the prompt kept growing. The docs list 33 hook event types, 46 built-in tool names and six permission modes, and Claude Code shipped between 20 and 29 releases in every month of 2026 through September. “Thin harness” therefore has two meanings that are easy to mix up: what the model reads, which keeps shrinking, and what you configure, which keeps growing.

The left half plays the two prompt cuts and the tool removals; each cut is relative to the prompt of its day, so read the bars as two separate events. The right half counts the surface as the docs listed it on October 6, 2026, with the releases for each month of the year.

The working form of this idea is an ablation. In his YC Startup School talk, as transcribed by a third party, Cherny described deleting “the entire system prompt” for each new model and bringing it back “line by line,” and advised users to delete their CLAUDE.md, skills and hooks every six months. Claude Code now ships /doctor prompt-audit, which looks for “instructions written for older models,” and its model guidance for Fable says “Skip the verification reminders: it verifies its own work with less prompting.”

Deletion is a harness change like any other, and harness changes can regress without anyone noticing. Anthropic's April 2026 postmortem traced degraded Claude Code quality to three harness-side changes while the API was unaffected. The default effort dropped from high to medium (“This was the wrong tradeoff.”). A change that cleared old thinking from idle sessions shipped with a bug that cleared it on every turn, passed review, tests and dogfooding, and took over a week to find. A verbosity line, “keep text between tool calls to ≤25 words,” cost 3% on one evaluation for both Opus 4.6 and 4.7. Removing scaffolding needs the same measurement as adding it, which Chapter 12 builds.

Turning a failure into a permanent fix

Mitchell Hashimoto's “My AI Adoption Journey” names the operator's core habit. When an agent makes a mistake, “you take the time to engineer a solution such that the agent never makes that mistake again.” He works in two forms: better implicit prompting, which means AGENTS.md lines, and “actual, programmed tools” such as screenshot scripts and filtered test runs. His linked example, Ghostty's src/inspector/AGENTS.md, is four bullet rules, each written after an observed bad behavior, among them a build flag (-Demit-macos-app=false) and the line “There are no unit tests in this package.” He reports that “it almost completely resolved them all.”

The habit scales past one person. Tobi Lütke reports that River, the agent that authors one in eight merged Shopify pull requests, raised its merge rate from 36% to 77% over two months with no retraining or model switch. People watched where it got stuck and wrote down what it should have known. OpenAI's harness post treats each struggle the same way, as a signal that asks “what capability is missing, and how do we make it both legible and enforceable for the agent?”

A fix can live at different heights. Anthropic's steering post lists seven mechanisms in Claude Code (CLAUDE.md, rules, skills, subagents, hooks, output styles and system-prompt appends) and sums up the choice: “Each method trades context cost against authority.” The memory docs call CLAUDE.md “context, not enforced configuration” and point to a PreToolUse hook when something must be blocked. The permission-mode docs say that a boundary stated in conversation, such as “don't push,” “can be lost if context compaction removes the message,” and that a hard guarantee needs a deny rule.

Hover over a rung, or tab to it, to see how it fails, with an example. Start each fix at the lowest rung that holds, and climb one rung when the same failure comes back.

Take one failure up the ladder: an agent force-pushes over a teammate's commits. A chat instruction can vanish with compaction and an AGENTS.md line is advisory. A permission rule such as Bash(git push --force *) matches command text, and the permission docs say a rule on git push misses git -C . push and git 'push'. A hook can parse the command instead: the script in the first tab handles git's global options, quoting, command chains, sh -c strings, bundled short flags and the +branch refspec, which forces a push without any flag. Claude Code and Codex both pass the hook the command as tool_input.command and treat exit code 2 as a block whose stderr the model reads, so one script serves both. The last tab is the same rule in the builder harness you will write in Chapter 2.

One rule, five places: no force pushes
.claude/hooks/no_force_push.py
python
#!/usr/bin/env python3
"""PreToolUse hook: refuse git force pushes however they are spelled.

Claude Code and Codex both send {"tool_input": {"command": "..."}} on stdin for
Bash calls and treat exit code 2 as a block, showing stderr to the model.
Not exhaustive (aliases, scripts, eval); the server-side setting is the backstop.
"""
import json
import shlex
import sys

FORCE = {"-f", "--force", "--force-with-lease", "--force-if-includes"}
WRAPPERS = {"command", "env", "exec", "nice", "nohup", "time"}
BREAKS = {"&&", "||", ";", "|", "&", "|&", "(", ")", "{", "}"}


def commands(line: str) -> list[list[str]]:
    """Split a shell command into simple commands, including sh -c strings."""
    out: list[list[str]] = []
    for physical in line.splitlines():
        lex = shlex.shlex(physical, posix=True, punctuation_chars=True)
        lex.whitespace_split = True
        cur: list[str] = []
        for tok in lex:
            if tok in BREAKS:
                out.append(cur)
                cur = []
            else:
                cur.append(tok)
        out.append(cur)
    nested = [c for argv in out
              if len(argv) > 2 and argv[0].split("/")[-1] in ("sh", "bash", "zsh") and argv[1] in ("-c", "-lc")
              for c in commands(argv[2])]
    return [argv for argv in out if argv] + nested


def is_force_push(argv: list[str]) -> bool:
    while argv and (argv[0] in WRAPPERS or ("=" in argv[0] and not argv[0].startswith("-"))):
        argv = argv[1:]  # drop wrappers and VAR=value prefixes
    if not argv or argv[0].split("/")[-1] != "git":
        return False
    i = 1
    while i < len(argv) and argv[i].startswith("-"):  # git -C dir, -c key=value, --no-pager
        i += 2 if argv[i] in ("-C", "-c") else 1
    if i >= len(argv) or argv[i] != "push":
        return False
    for arg in argv[i + 1:]:
        if arg in FORCE or arg.startswith("--force-with-lease=") or arg.startswith("+"):
            return True
        if arg.startswith("-") and not arg.startswith("--") and "f" in arg[1:]:
            return True  # bundled short flags such as -fu
    return False


def main() -> None:
    command = json.load(sys.stdin).get("tool_input", {}).get("command", "")
    try:
        found = any(is_force_push(argv) for argv in commands(command))
    except ValueError:  # unbalanced quotes: refuse anything that looks like a push
        found = "git" in command and "push" in command
    if found:
        print("Blocked: force pushes rewrite shared history. Push to a new branch "
              "and open a pull request instead.", file=sys.stderr)
        sys.exit(2)


if __name__ == "__main__":
    main()

Make the hook executable with chmod +x. Codex's prefix rules have a blind spot of their own. A pattern must match an exact prefix, so the rule refuses git push --force origin main and lets git push origin main --force through; the not_match entry records that, and codex execpolicy check tests any command against the file. Codex also asks you to review a new hook with /hooks and records your trust against the hook's hash. The top rung for this failure sits on the server. With receive.denyNonFastForwards set on the shared repository, Git refuses a non-fast-forward update “even if that push is forced,” whichever harness, model or person sent it.

Where a “don't” belongs
A rule that must hold every time needs a rung that enforces it: a deny rule, a hook, a test, or a change to the repository or the server. AGENTS.md is for what the agent cannot find out on its own, like Ghostty's build flag. Each line costs context on every turn, and Claude Code's memory docs target “under 200 lines per CLAUDE.md file” because “Longer files consume more context and reduce adherence.”

The ledger is a JSON Lines file in the repository with one entry per failure: when it happened, which harness and model, the symptom, the path to the transcript, the rung the fix went to, the file it changed, and a short exact phrase from the fix so the report can find it. Copy the transcript next to the ledger so it survives the harness's own cleanup. When a fixed failure comes back, add an entry that names the original in recurs.

failures.jsonl
json
{"id": "F-001", "date": "2026-09-29", "harness": "claude-code 2.1.291, claude-sonnet-5-5", "symptom": "ran the whole test suite after a one-line change", "transcript": "ledger/F-001.jsonl", "layer": "agents_md", "file": "AGENTS.md", "fix": "Run only the tests for the package you changed"}
{"id": "F-002", "date": "2026-09-30", "harness": "claude-code 2.1.291, claude-opus-5-5", "symptom": "force-pushed over a teammate's commits on a shared branch", "transcript": "ledger/F-002.jsonl", "layer": "hook", "file": ".claude/hooks/no_force_push.py", "fix": "def is_force_push"}
{"id": "F-003", "date": "2026-10-01", "harness": "claude-code 2.1.291, claude-opus-5-5", "symptom": "hand-edited the generated API client in src/gen/", "transcript": "ledger/F-003.jsonl", "layer": "agents_md", "file": "AGENTS.md", "fix": "Never edit files under src/gen/"}
{"id": "F-004", "date": "2026-10-03", "harness": "claude-code 2.1.291, claude-opus-5-5", "symptom": "edited src/gen/client.ts again despite the AGENTS.md line", "transcript": "ledger/F-004.jsonl", "recurs": "F-003", "layer": "hook", "file": ".claude/settings.json", "fix": "Edit(/src/gen/**)"}
{"id": "F-005", "date": "2026-10-04", "harness": "claude-code 2.1.291, claude-sonnet-5-5", "symptom": "imported a server-only module into the browser bundle", "transcript": "ledger/F-005.jsonl", "layer": "lint", "file": "eslint.config.js", "fix": "no-restricted-imports"}
{"id": "F-006", "date": "2026-10-05", "harness": "codex 0.160.1, gpt-6.1-sol", "symptom": "could not find how to boot the app and guessed a port", "transcript": "ledger/F-006.jsonl", "layer": "none"}

ledger.py prints fixes and recurrences per rung, each recurrence (with a suggestion to climb when the new fix sits no higher than the one that failed), the open failures, and the AGENTS.md lines that no logged failure justifies.

ledger.py
python
#!/usr/bin/env python3
"""Report on the failure ledger: fixes per rung, recurrences, and the
AGENTS.md lines that no logged failure justifies.

    python ledger.py failures.jsonl --agents-md AGENTS.md

Each entry names the rung its fix went to ("layer"), the file it changed and a
short exact phrase from the fix ("fix"); a recurrence names the original in "recurs".
"""
import argparse
import json
from collections import Counter
from pathlib import Path

# The fix ladder, lowest rung first. "hook" covers permission rules as well
# (Claude Code deny and ask rules, Codex prefix_rule).
LADDER = ["chat", "agents_md", "skill", "hook", "lint", "test", "structure"]


def load(path: Path) -> list[dict]:
    rows = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]
    for r in rows:
        if r.get("layer", "none") not in LADDER + ["none"]:
            raise SystemExit(f"{r['id']}: unknown layer {r['layer']!r}")
    return rows


def main() -> None:
    ap = argparse.ArgumentParser()
    ap.add_argument("ledger", type=Path)
    ap.add_argument("--agents-md", type=Path)
    args = ap.parse_args()
    rows = load(args.ledger)
    by_id = {r["id"]: r for r in rows}

    fixed = Counter(r["layer"] for r in rows if r.get("layer", "none") != "none")
    # A recurrence counts against the rung of the fix that did not hold.
    again = Counter(by_id[r["recurs"]]["layer"] for r in rows if r.get("recurs") in by_id)
    print(f"{'rung':<11}{'fixes':>6}{'recurred':>10}")
    for rung in LADDER:
        if fixed[rung] or again[rung]:
            print(f"{rung:<11}{fixed[rung]:>6}{again[rung]:>10}")

    print()
    for r in rows:
        first = by_id.get(r.get("recurs", ""))
        now = r.get("layer", "none")
        if first:
            note = f"{first['id']} ({first['layer']}) came back as {r['id']}, "
            note += f"now fixed at {now}" if now != "none" else "still open"
            if now in LADDER and LADDER.index(now) <= LADDER.index(first["layer"]) < len(LADDER) - 1:
                note += f"; climb to {LADDER[LADDER.index(first['layer']) + 1]} or higher"
            print(note)
        elif now == "none":
            print(f"open: {r['id']} {r['symptom']} ({r['transcript']})")

    if args.agents_md:
        text = args.agents_md.read_text()
        mine = [r for r in rows if r.get("file") == args.agents_md.name and r.get("fix")]
        rules = [ln.strip() for ln in text.splitlines() if ln.lstrip().startswith(("- ", "* "))]
        orphans = [ln for ln in rules if not any(r["fix"].lower() in ln.lower() for r in mine)]
        print()
        print(f"{args.agents_md}: {len(rules)} rules, {len(orphans)} with no logged failure")
        for ln in orphans:
            print(f"  {ln}")
        for r in mine:
            if r["fix"].lower() not in text.lower():
                print(f"stale: the fix for {r['id']} is no longer in {args.agents_md}")


if __name__ == "__main__":
    main()

On the six entries above, the report shows the climb from F-003 to F-004 and F-006 still open. An AGENTS.md line with no failure behind it deserves a second look: some carry knowledge the agent could never discover, and the rest are candidates for the deletions of section 7. Chapter 3 uses the ledger to decide what AGENTS.md keeps, Chapter 6 turns repeated “done but not done” entries into gates, and Chapter 12 turns entries into evaluation tasks, the step our post on turning failures into evals walks through.

Reading a harness claim

Every benchmark number in this book carries a triple: the model, with its effort setting when the source gives one; the harness, with its version; and the benchmark version. Terminal-Bench 2.0, 2.1, 3.0 and 4.0 are different benchmarks, and so are SWE-bench Verified, SWE-bench Pro and SWE-Bench Pro V2. The same Opus 4.6 and Claude Code pair scores 58.0% on Terminal-Bench 2.0 and 70.1% on 2.1, and boards change between snapshots: on October 6, 2026 the Terminal-Bench 2.0 board lists Opus 4.6 in ten harnesses, one more than in August.

Past the triple, five things decide how far a number carries:

  • Trials and intervals. The Terminal-Bench 2.0 paper reports 95% intervals of two to three points, so a smaller gap needs more trials before it means anything.
  • Who ran it. Verified rows and author-run tables carry more weight than vendor results or self-submitted entries.
  • Resources and configuration. In Anthropic's infrastructure-noise study, resources alone moved Terminal-Bench 2.0 by 6 points, and “leaderboard differences below 3 percentage points deserve skepticism until the eval configuration is documented and matched.”
  • Leakage. Whether the agent could reach the tests, the answers, or a verifier that trusts its output, as in the integrity cases in section 3.
  • Cost. Tokens and dollars per solved task, which move even when success does not; this is where Reflection's “Pick the harness for the bill” applies.

The research behind this book grades every fact it uses, and the text tells you which grade a claim rests on.

GradeWhat it meansHow the text writes itExample in this chapter
PrimaryRead at the source: a paper, docs page, changelog, data file or the company's own postStated plainly, with a linkThe Terminal-Bench 2.0 board rows; Claude Code's hook and permission docs
Primary, abstract onlyOnly the paper's abstract was readUses only what the abstract statesLewis's 28% to 49%
SecondarySeen only in someone else's reportAttributed in the sentence: “as reported by”Cherny's “around half the system prompt”, via the Pragmatic Engineer
Vendor or company self-reportPrimary as to what the company said, not as to whether it is true“X reports”, never presented as an independent measurementCline's 6.34% to 0.62% A/B
UnconfirmedSeen only in a search answer or on a page that could not be readLeft out, or named as circulating and unconfirmedNone used in this chapter

Grade the chapter's opening claim this way. “Opus 4.6 spans 18.4 points across nine harnesses on Terminal-Bench 2.0” has a full triple and a snapshot date, and it is primary as to what the board shows. But seven of its nine rows are unverified, the board does not control resources or configuration (as an independent analysis of it noted), and Meerkat caught a trace from the top row printing “PASS” to fool a verifier. What survives is narrower. On verified rows the spread is 4.9 points. In the controlled studies of section 2, harness changes moved fixed models by about 7 to 21 points, and much further in the extreme cases of a poor edit format or a tight context budget. For strong models on SWE-style tasks, several studies find almost no effect on success and a clear one on cost.

When a vendor says its harness is better, ask for the triple, the trials, who ran it, the resources and the cost per solved task, then run both harnesses on your own tasks with several trials each, the rig Chapter 12 builds. Our book Evaluating Agentic AI covers pass^k and paired comparisons in depth.

Chapter 1 in one page

Key takeaways
8 items
  • 1Holding the model fixed, harness changes moved LangChain's agent from 52.8% to 66.5% on Terminal-Bench 2.0, AHE's from 69.7% to 77.0% on Terminal-Bench 2, and Lewis's fail-to-pass fraction from 28% to 49%.
  • 2Opus 4.6's 18.4-point spread across nine Terminal-Bench 2.0 harnesses is 4.9 points on verified rows. Check the verified flag, the snapshot date and the integrity notes before you quote a board.
  • 3The effect is largest for weaker models, tight context budgets and long terminal tasks (context management was worth 35.7 points at 32k and 2.7 at 128k in one study) and small for strong models on SWE-style tasks, where the harness mostly moves cost.
  • 4The harness is everything that is not the model. Keep the builder harness (code you own, like hx) apart from the user harness (the configuration and repository you add to a harness someone else wrote).
  • 5The earliest documented use of “harness engineering” is Dex Horthy's post of November 4, 2025; HumanLayer, Osmani, arXiv 2609.00006 and Wikipedia credit others.
  • 6Every component encodes an assumption about the model. Claude Code removed over 80% of its system prompt for Opus 5 and Fable 5 with no measurable loss, and Anthropic dropped context resets once Opus 4.5 stopped needing them.
  • 7Log every failure with its transcript, fix it at the lowest rung that holds, and climb when it recurs. A rule that must always hold belongs in a deny rule, a hook, a test or on the server.
  • 8Read every number by its triple (model and effort, harness and version, benchmark version) and by its grade: primary, secondary or self-report.

What to do on Monday. Add failures.jsonl to the repository your agent works in and log the next three failures, each with its transcript and the rung you fixed it at. Run capture_requests.py and first_turn_audit.py once against the harness you use most and write down its first-request total. Then run ledger.py with --agents-md AGENTS.md and delete one line that no failure justifies.

Chapter 2 · The loop

The Loop, One Turn at a Time

The agent loop: context, model, tool call and observation, with stop rules

A working definition of a harness, one turn traced through Codex and Claude Code, and a small Python harness with stop rules, a dollar cap and an event log that can rebuild every request the model saw.

Every coding agent runs the same cycle: the harness sends the model a request, reads back a tool call, runs it, appends the result and decides whether to go again. Claude Code, Codex and a 100-line research script differ in what they put into each request and what they let out of it. This chapter takes the cycle apart one turn at a time and then builds it.

The minimal version scores well. On the Terminal-Bench 3.0 board the top entry is Claude Opus 5 in mini-SWE-agent (“some 100 lines of python”, per its README) at 42.7%, ahead of GPT-5.6 Sol in Codex (34.59%) and Claude Fable 5 in Claude Code (34.05%). The Terminal-Bench 4.0 board is led by vendor harnesses as of October 6, 2026 (Opus 5.5 in Claude Code at 64.85%), so a small loop is a strong baseline that every added component has to beat on your own tasks.

You will finish with hx/loop.py and hx/events.py, the core of the Python harness every later chapter extends, and with the vocabulary the rest of the book uses: turn, observation, stop rule, budget and event log.

A harness in one sentence

LangChain's Vivek Trivedy drew the boundary in one line: “If you're not the model, you're the harness.” His longer definition in The Anatomy of an Agent Harness (March 2026) is “every piece of code, configuration, and execution logic that isn't the model itself”, hence “Agent = Model + Harness”. Claude Code's glossary says the same of its own product: “Claude Code is the harness; Claude is the model inside it.”

Those definitions say where the harness ends. To build one, this book uses a definition of what it does:

A harness is the program that decides what the model sees on every request and what happens to everything the model asks for.

The first half is context: which instructions, tool definitions, files and earlier results go into the request, in what order. The second half is action: which calls run, where, under what limits, and what comes back. Anthropic's Lance Martin lists the same parts as “the loop, tools, context management, and guardrails” (Harnessing Claude's intelligence). The cycle joining the halves is Simon Willison's definition of an agent, “Agents run tools in a loop to achieve a goal” (Agentic Engineering Patterns), which Claude Code's docs split into three blended phases: gather context, take action, verify results (How Claude Code works).

Birgitta Böckeler (martinfowler.com) separates the builder harness the vendor ships (loop, tools, compaction) from the user harness you wrap around it (instruction files, skills, hooks, checks), and calls engineering the user harness “a specific form of context engineering”. This chapter builds a builder harness from nothing, so the user-harness chapters (3, 6 and 10) have something concrete to plug into.

The core is small everywhere. A July 2026 audit of eleven production harnesses, Claude Code, Codex CLI and Pi among them (arXiv 2609.00006), found that none of their roughly 4 million lines imports a general agent framework, and it ends with a 90-line minimum harness. Our post on what 11 production coding agents share walks through the audit.

Anatomy of a turn

Sources use “turn” for different things. DeepSeek Harness calls one model request plus its tool calls a step, and a turn zero or more steps (architecture doc); for the Codex App Server a turn is one unit of agent work started by user input (Unlocking the Codex harness). The Claude Agent SDK's maxTurns counts “tool-use round trips”, and this book follows it: a turn is one model request and the tool calls in its response, so max_turns=30 means at most 30 model calls.

Michael Bolin's Unrolling the Codex agent loop (OpenAI, January 2026) documents the order in which Codex assembles a request:

  1. instructions: the file named by model_instructions_file if set, otherwise the model's bundled base prompt, such as gpt-5.2-codex_prompt.md.
  2. tools: shell, update_plan, web_search and any MCP tools.
  3. input items, in order: a developer message describing permissions (built from snippets such as workspace_write.md and on_request.md), an optional developer_instructions message, a user message with the aggregated instructions (AGENTS.override.md or AGENTS.md from CODEX_HOME, then each folder from the project root down to the working directory, up to 32 KiB, then skills metadata), a user message holding <environment_context> with the working directory and shell, and last the user's message.

The figure walks one turn in that order: the request on the left fills one layer per step, each tagged with who supplies it, and the boxes on the right show who acts.

It opens on the last step, where the request, the response and the tool output together become the prefix of the next request; press NEXT to start from the first layer, or PLAY to watch. Of the nine items, the harness writes five, you write two, the repository supplies one and the model writes one.

Three details are decisions for your own harness. Stable parts come first, because the cache hits only on an exact prefix. The sandbox “applies only to the Codex-provided shell tool”, so MCP tools “are responsible for enforcing their own guardrails”. And a mid-session change of sandbox, approval mode or working directory is appended as a new message instead of edited into the old one.

Claude Code's request has the same shape. Its prompt caching page lists the system prompt and tool definitions, then project context (CLAUDE.md, auto memory, unscoped rules), then the conversation, and Thariq Shihipar gives the reason in Prompt caching is everything: “At Claude Code, we build our entire harness around prompt caching.”

In one turn the harness assembles the request, sends it, parses the response, decides whether each tool call may run, runs it, turns the output into an observation, appends it, logs everything and checks the stop rules, while the model's whole part is sampling the response.

Stateless requests, growing transcripts

Codex does not use the Responses API's previous_response_id. Bolin writes that keeping requests “fully stateless” supports Zero Data Retention, with encrypted reasoning from earlier turns decrypted on the server. Every request carries the whole conversation, as in hx, where state.messages is the session and every call sends all of it.

Each request therefore contains the previous one as an exact prefix. Bolin spells out the cost: the loop is quadratic in the JSON sent, and with prompt-cache hits sampling becomes “linear rather than quadratic”, as long as the prefix is exact (“note how the old prompt is an exact prefix of the new prompt.”). The figure runs the task from Anthropic's What a task costs on Opus 5.5, whose context grows from 20K to 120K tokens over 40 turns, drawn here as a straight line.

Drag the slider to change how far the task has run. Each bar is one request, the previous one plus about 2,564 new tokens, and the area on the right is everything sent so far. PREFIX CACHE shows the same traffic with an exact-prefix cache: each request after the first reads the whole previous request from cache (hatched) and processes only the new part.

At 40 turns the run has sent 2.8 million input tokens to build a 120,000-token transcript, which Anthropic's post prices on Opus 5.5 at $11.20 uncached and about $1.62 at a 90% cache hit rate: “A turn costs more than the tokens it adds, because it resends everything before it.” A perfect prefix cache reaches about 96% (the post's $0.99 case), since only 120,000 tokens are ever processed new. The same growth in 25 turns sends about 1.75 million tokens ($1.02), so turn count is a cost lever before caching enters. Chapter 4 does the economics.

Input dominates agent traffic. Manus reported an input-to-output ratio of about 100:1 (context engineering post, July 2025), and Anthropic reports Claude Code's ratio moving from 189:1 to 324:1 between March and September 2026 as context per request grew 2.6x (Opus 5.5 announcement).

An edit to an early part of the request breaks the prefix for everything after it. Bolin lists what has broken it in Codex: changing tools, model, sandbox configuration, approval mode or working directory mid-conversation, an MCP bug that listed tools in inconsistent order, and notifications/tools/list_changed messages. Both vendors append instead; Anthropic's API takes mid-conversation messages with role system because “editing the top-level system field changes the very beginning of the prompt and invalidates the cache for everything that follows” (docs). In hx, before_request may rewrite history, as Chapter 4's reduction ladder does, at the price of a cache miss after the change.

Stateless is a choice

OpenAI's WebSocket mode keeps a connection-scoped, in-memory copy of the previous response state keyed by previous_response_id, and OpenAI reports agent loops 40% faster end to end with it (Speeding up agentic workflows with WebSockets). Codex gives that up for Zero Data Retention and a request any process can rebuild; hx stays stateless for the second reason.

The smallest loop that scores

mini-SWE-agent, from the SWE-agent team, asks in its README: “What if our agent was 100x simpler, and still worked nearly as well?” Its agent class is “some 100 lines of python”: bash as the only tool, a fully linear history, each action in its own subprocess (with a 30-second default timeout in version 2), and a claimed score above 74% on SWE-bench Verified. Version 2's agent file was 190 lines on 2026-10-06, 171 of them code. It uses native tool calling with one bash tool, ends when the model runs echo COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT, and sets step_limit 0 (off), cost_limit 3.0 and max_consecutive_format_errors 3 in AgentConfig.

Besides its Terminal-Bench 3.0 lead, Datacurve's DeepSWE ran ten SWE-Bench Pro tasks through mini-swe-agent and each vendor's harness: Opus 4.7 passed 50% in mini against 40% in Claude Code, Gemini 3.1 Pro 40% against 20% in Gemini CLI, and GPT-5.5 tied at 40% with Codex CLI. The authors call the sample small and say part of the gap is likely prompt tuning.

Mario Zechner's Pi has four tools (read, write, edit, bash), and “pi's system prompt and tool definitions together come in below 1000 tokens.” In November 2025 it had no to-do list, plan mode, MCP, background bash or subagents, and no permission prompts. Zechner concluded that “these four tools are all you need for an effective coding agent”, following a rule he states as “if I don't need it, it won't be built.”

Terminus 2, the Terminal-Bench authors' reference agent, is “a minimal, unopinionated agent”: one headless tmux terminal driven by keystrokes, plus summarization when the context fills. Its episode count was uncorrelated with success (r = −0.028, p = 0.916; arXiv 2601.11868), so longer runs were no more likely to succeed.

Controlled comparisons agree. Scale and Reflection ran GLM-5.3, Kimi K3 and Inkling through mini-swe-agent, Pi and OpenCode on 642 SWE-Bench Pro V2 tasks under a locked budget and found the strongest models on par (Inkling did worse on Pi); Reflection's advice is “Pick the harness for the bill, not the score” (Scale Labs). METR found neither Claude Code nor Codex beat its simple scaffolds on time horizon (Claude Code beat ReAct in 50.7% of bootstrap samples for Opus 4.5, Codex beat Triframe in 14.5% for GPT-5; METR note).

Claude Code's Boris Cherny makes the same point from inside a large harness: “when the model is so good, the simple thing usually works” (transcript), and in July 2026 Anthropic reported that it “removed over 80% of Claude Code's system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations” (The new rules of context engineering).

Minimal loops lose ground in specific places. On the Terminal-Bench 2.0 leaderboard's verified rows, GPT-5 scores 49.6% in Codex CLI and 33.9% in mini-SWE-agent, and a September 2026 study of ten Qwen and DeepSeek models found that complex harnesses help less on SWE-style repair as models improve but still help stronger models on open-ended repository tasks (arXiv 2609.32459, abstract). Local models show the same dependence on fit (model plus loop fit beats the biggest context). Section 10 lists where the minimal loop breaks.

Stop rules and budgets

A loop needs two kinds of exit. The natural one is the model saying it is done: in hx a response with no tool call, in mini-SWE-agent the sentinel command. Everything else is a stop rule, a condition the harness checks between requests that ends the run or sends it back to work.

Natural exits fail both ways. Models stop early: LangChain's most common Terminal-Bench 2.0 failure was that “the agent wrote a solution, re-read its own code, confirmed it looks ok, and stopped” (Improving Deep Agents with harness engineering), and Claude Code's best practices warn that “Claude stops when the work looks done.” Models also fail to stop, and those runs cost the most: Scale and Reflection found failed runs used more steps and tokens than solved ones at the median in 8 of 9 model-harness combinations. The rules in use, with their published defaults:

RuleWhat it catchesDefaults in the sourcesIn hx
Turn caploops that never convergeSymphony agent.max_turns 20; Agent SDK maxTurns; mini-SWE-agent step_limit (0, off)max_turns in the core
Dollar capruns whose context or output balloonsmini-SWE-agent cost_limit 3.0; Agent SDK maxBudgetUsd; Claude Code --max-budget-usdmax_usd in the core, priced from usage
Completion checkstopping at the first plausible answerClaude Code Stop hook, overridden after 8 consecutive blocks; /goal; LangChain PreCompletionChecklistMiddlewarebefore_stop (Chapter 6)
Repeat detectorthe same failing call, again and againreminder after 5 identical calls, stop after 8 identical failing calls (arXiv 2609.20804); LangChain LoopDetectionMiddlewaremiddleware (Chapter 6)
Malformed-call limitcalls the harness cannot runmini-SWE-agent max_consecutive_format_errors 3; Cline aborts after three consecutive mistakesMalformedLimit in section 6
Stall timeouta command or session that never returnsSymphony stall_timeout_ms 300000 (5 minutes); Claude Code BASH_DEFAULT_TIMEOUT_MS 120000the bash tool's 60-second timeout

The figure runs five failing sessions against five of these rules, starting with the two caps the hx core has switched on.

Toggle the rules and watch the outcome column. Caps end every runaway run, late, and say only that it was long; detectors end runs earlier and name a reason, such as an eighth identical failing call. The run polling a stuck job makes successful calls, so the repeat detector only reminds it. The run that says it is done without testing gets past every cap, and only a completion check that can tell tested from untested catches it (Chapter 6).

The repeat detector's thresholds come from the harness study in arXiv 2609.20804, the turn cap and stall timeout from OpenAI's Symphony spec. Claude Code's Stop hook can block a stop and hand the model a reason; so that a broken hook cannot loop forever, Claude Code overrides it after 8 consecutive blocks, resetting the count when Claude calls a tool (CLAUDE_CODE_STOP_HOOK_BLOCK_CAP raises the cap; hooks reference). /goal wraps the same mechanism, with a small, fast model judging the transcript “not yet met”, “met” or “impossible” after each turn (goal docs).

A stop rule's firing rate is also a health metric. Cline's task.mistake_limit_reached (three consecutive mistakes) fell from 6.34% to 0.62% of tasks when its VS Code extension (11M+ users) moved to a new SDK harness, a drop Cline attributes to the old design around XML tool calls parsed from text (company report).

hx checks the caps before each request, so a run overshoots max_usd by at most one response. A cap the model cannot see can only cut it off; LangChain added time-budget warnings that tell the agent to wrap up, and Anthropic's task budgets (output_config.task_budget, beta header task-budgets-2026-03-13; taskBudget in the Agent SDK, still alpha) tell the model how much of its token budget remains.

The same rules in Claude Code and Codex
  • Claude Code print mode takes both caps, claude -p --max-turns 30 --max-budget-usd 1.00, and both flags work in print mode only.
  • In the Agent SDK, a run stopped by maxTurns or maxBudgetUsd ends with a result whose subtype is error_max_turns or error_max_budget_usd.
  • A completion check is a Stop hook (Chapter 6 writes one), or /goal when a judged condition is enough.
  • In Codex, App Server goals carry a tokenBudget, and codex exec --json reports token usage on every turn.completed event for a wrapper script to cap.

Build: a minimal harness

Here is the core, written to the contract every later chapter imports; it needs Python 3.11+ and the anthropic package. The file is about 140 lines with comments, smaller than mini-SWE-agent's 190-line agent file, and run() is about 80. The first half defines the shared types.

hx/loop.py (part 1 of 2): types and prices
python
"""hx/loop.py: the core loop of the hx harness (Harness Engineering, Chapter 2).

pip install anthropic   (Python 3.11+; the client reads ANTHROPIC_API_KEY)
"""
import json
from dataclasses import dataclass, field
from typing import Callable

import anthropic

from hx.events import EventLog, build_request, fingerprint

# USD per million tokens: input, output, cache read, 5-minute cache write.
# Anthropic pricing page, accessed 2026-10-06. Add a row before using another model.
PRICES = {
    "claude-sonnet-5-5": (2.00, 10.00, 0.20, 2.50),
    "claude-opus-5-5": (4.00, 20.00, 0.20, 5.00),
    "claude-fable-5-1": (10.00, 50.00, 0.25, 12.50),
    "claude-haiku-4-5-20251001": (1.00, 5.00, 0.10, 1.25),
}
MAX_TOKENS = 16_000  # output cap per response


@dataclass
class Tool:
    name: str
    description: str
    input_schema: dict              # JSON Schema for the tool input
    run: Callable[[dict], str]      # returns the observation text the model reads


@dataclass
class State:
    task: str
    system: str
    messages: list[dict]            # Anthropic Messages API format
    turn: int = 0
    usage: dict = field(default_factory=lambda: {"input": 0, "output": 0, "cache_read": 0, "cache_write": 0})
    cost_usd: float = 0.0
    stop_reason: str | None = None  # "done" | "max_turns" | "budget" | "gate" | "error"


class Middleware:                   # no-op defaults; subclasses override what they need
    def before_request(self, state: State) -> None: ...
    def before_tool(self, state: State, name: str, args: dict) -> str | None: return None
    def after_tool(self, state: State, name: str, args: dict, result: str) -> str: return result
    def before_stop(self, state: State, final_text: str) -> str | None: return None


def price(model: str, u) -> float:
    p_in, p_out, p_read, p_write = PRICES[model]
    return (u.input_tokens * p_in + u.output_tokens * p_out
            + (u.cache_read_input_tokens or 0) * p_read
            + (u.cache_creation_input_tokens or 0) * p_write) / 1_000_000

PRICES holds Anthropic's list prices as of October 6, 2026 (pricing page): Sonnet 5.5 costs $2 per million input tokens, $10 per million output, $0.20 per million cache reads and $2.50 per million five-minute cache writes. Every response is priced from its usage field, cache columns included, so the dollar cap sees the discount once Chapter 4 adds caching.

hx/loop.py (part 2 of 2): run()
python
def run(task: str, tools: list[Tool], *, system: str = "", model: str = "claude-sonnet-5-5",
        max_turns: int = 50, max_usd: float = 5.0, middleware: list[Middleware] | None = None,
        log_path: str | None = None) -> State:
    if model not in PRICES:
        raise ValueError(f"no price for {model}; add it to PRICES so the dollar cap works")
    client, mws, log = anthropic.Anthropic(), middleware or [], EventLog(log_path) if log_path else None
    by_name = {t.name: t for t in tools}
    specs = [{"name": t.name, "description": t.description, "input_schema": t.input_schema}
             for t in tools]
    state = State(task=task, system=system, messages=[])

    def emit(kind: str, data: dict) -> None:
        if log:
            log.append(kind, state.turn, data)

    def add(message: dict) -> None:
        state.messages.append(message)
        emit("message", message)

    def call_tool(name: str, args: dict) -> tuple[str, bool]:
        verdicts = [m.before_tool(state, name, args) for m in mws]
        denial = next((v for v in verdicts if v is not None), None)     # first denial wins
        if denial is not None:
            return denial, True
        try:
            out = by_name[name].run(args) if name in by_name else f"error: no tool named {name!r}"
        except Exception as e:                       # a tool crash becomes an observation
            out = f"error: {type(e).__name__}: {e}"
        for m in mws:
            out = m.after_tool(state, name, args, out)
        return out, False

    emit("start", {"model": model, "max_tokens": MAX_TOKENS, "system": system, "tools": specs})
    add({"role": "user", "content": task})
    while state.stop_reason is None:
        if state.turn >= max_turns or state.cost_usd >= max_usd:
            state.stop_reason = "max_turns" if state.turn >= max_turns else "budget"
            break
        state.turn += 1
        before = json.dumps([state.system, state.messages])
        for m in mws:
            m.before_request(state)
        if json.dumps([state.system, state.messages]) != before:       # history was rewritten
            emit("rewrite", {"system": state.system, "messages": state.messages})
        if state.stop_reason:                                           # or the run was ended
            break
        request = build_request(model, MAX_TOKENS, state.system, specs, state.messages)
        emit("request", {"sha256": fingerprint(request), "messages": len(state.messages)})
        try:
            resp = client.messages.create(**request)
        except anthropic.APIError as e:
            emit("error", {"error": repr(e)})
            state.stop_reason = "error"
            break
        u = resp.usage
        state.usage["input"] += u.input_tokens
        state.usage["output"] += u.output_tokens
        state.usage["cache_read"] += u.cache_read_input_tokens or 0
        state.usage["cache_write"] += u.cache_creation_input_tokens or 0
        state.cost_usd += price(model, u)
        emit("response", {"stop_reason": resp.stop_reason, "usage": u.to_dict(mode="json"),
                          "cost_usd": round(state.cost_usd, 6)})
        add({"role": "assistant",
             "content": [b.to_dict(mode="json", exclude_none=True) for b in resp.content]})
        calls = [b for b in resp.content if b.type == "tool_use"]
        if not calls:                                 # a reply with no tool call: may it stop?
            final = "".join(b.text for b in resp.content if b.type == "text")
            nudges = [m.before_stop(state, final) for m in mws]
            nudge = next((n for n in nudges if n is not None), None)
            if nudge is None:
                state.stop_reason = state.stop_reason or "done"
            else:
                add({"role": "user", "content": nudge})
            continue
        results = []
        for c in calls:
            args = dict(c.input)
            out, denied = call_tool(c.name, args)
            emit("tool", {"id": c.id, "name": c.name, "args": args, "result": out, "denied": denied})
            results.append({"type": "tool_result", "tool_use_id": c.id, "content": out})
        add({"role": "user", "content": results})
    emit("stop", {"stop_reason": state.stop_reason, "cost_usd": round(state.cost_usd, 6)})
    return state

The decisions in run(), in the order a request meets them:

  • Caps first. Both caps are checked before every request, and a middleware ends the run by setting state.stop_reason (gate for Chapter 6's gates, error for the limit below).
  • Only before_request rewrites history. Any change it makes is logged as a rewrite event, so the log can rebuild requests after compaction.
  • Messages stay plain dictionaries. Content blocks go through the SDK's to_dict, so middleware can edit them and the log can store them as JSON.
  • Tool failures become observations. An unknown tool or a tool exception becomes an error: string the model reads; an API error ends the run as error.
  • Every hook runs in list order, and the first answer wins. The first string from before_tool denies the call and becomes its result; the first from before_stop goes back to the model. before_tool, after_tool and before_stop mirror Claude Code's PreToolUse, PostToolUse and Stop hooks.
  • Three gaps wait for later chapters. No cache_control, so every turn pays the full input price (Chapter 4); no sandbox (Chapter 7); no output budget beyond a character cut (Chapter 5).

Four hooks cover what production harnesses intercept: DeepSeek Harness has six interception points, LangChain's agent middleware six hooks from before_agent to after_agent (middleware post); Claude Code has 33 hook events plus mods, “hooks that ship inside plugins” (mods post). A check written as hx middleware maps onto each of them.

The log calls go to hx/events.py, which stores the parts of each request and folds them back together on demand; section 7 explains why.

hx/events.py
python
"""hx/events.py: the hx session log (Harness Engineering, Chapter 2).

One JSON object per line, appended and never edited. Model-visible means logged:
every request run() sends can be rebuilt from the events logged before it.
"""
import hashlib
import json
import sys
import time
from pathlib import Path


class EventLog:
    def __init__(self, path: str):
        self.path = Path(path)
        self.path.parent.mkdir(parents=True, exist_ok=True)

    def append(self, kind: str, turn: int, data: dict) -> None:
        line = json.dumps({"ts": round(time.time(), 3), "kind": kind, "turn": turn, "data": data})
        with self.path.open("a", encoding="utf-8") as f:   # one write per event; a crash
            f.write(line + "\n")                            # can tear only the last line


def build_request(model: str, max_tokens: int, system: str, tools: list, messages: list) -> dict:
    """The request body, built the same way to send it and to rebuild it."""
    request = {"model": model, "max_tokens": max_tokens, "messages": messages}
    if system:
        request["system"] = system
    if tools:
        request["tools"] = tools
    return request


def fingerprint(request: dict) -> str:
    return hashlib.sha256(json.dumps(request, sort_keys=True).encode()).hexdigest()[:16]


def read_events(path: str) -> list[dict]:
    events = []
    for line in Path(path).read_text(encoding="utf-8").splitlines():
        try:
            events.append(json.loads(line))
        except json.JSONDecodeError:    # a line torn by a crash; everything before it is intact
            break
    return events


def requests(path: str):
    """Fold the log; yield (turn, logged fingerprint, rebuilt request) per request."""
    start, system, messages = {}, "", []
    for e in read_events(path):
        kind, data = e["kind"], e["data"]
        if kind == "start":
            start, system = data, data["system"]
        elif kind == "message":
            messages.append(data)
        elif kind == "rewrite":
            system, messages = data["system"], list(data["messages"])
        elif kind == "request":
            yield e["turn"], data["sha256"], build_request(
                start["model"], start["max_tokens"], system, start["tools"], list(messages))


def rebuild_request(path: str, turn: int) -> dict:
    """The exact request body run() sent on that turn."""
    for t, _, request in requests(path):
        if t == turn:
            return request
    raise KeyError(f"no request for turn {turn} in {path}")


def verify(path: str) -> None:
    """Rebuild every request, check its fingerprint, and flag turns that broke the prefix."""
    prev = None
    for turn, logged, req in requests(path):
        status = "ok" if fingerprint(req) == logged else "MISMATCH"
        same_prefix = prev is None or (req.get("system") == prev.get("system")
                                       and req["messages"][:len(prev["messages"])] == prev["messages"])
        note = "" if same_prefix else "  prefix changed (cache misses after the change)"
        print(f"turn {turn:>3}  {status}  {len(req['messages']):>3} messages{note}")
        prev = req


if __name__ == "__main__":
    verify(sys.argv[1])

The loop and the log share build_request, and fingerprint checks that sent and rebuilt requests match. Chapter 5 builds the full bash tool (hx/tools/bash.py); for now a minimal one lives in the example script, beside a middleware that ends the run after three unrunnable tool calls in a row, mini-SWE-agent's limit.

example_run.py
python
"""example_run.py: one bash tool, one middleware and one run of the hx loop (Chapter 2).

HX_WORKDIR=~/scratch/some-repo python example_run.py "Run pytest -q and fix what fails."
POSIX only. The tool runs commands as you: use a throwaway clone or a container.
"""
import contextlib
import os
import signal
import subprocess
import sys

from hx.loop import Middleware, State, Tool, run

WORKDIR = os.path.expanduser(os.environ.get("HX_WORKDIR", "."))
TIMEOUT_S = 60
KEEP = 8_000  # characters the model sees; Chapter 5's bash tool budgets output properly


def bash(args: dict) -> str:
    proc = subprocess.Popen(["bash", "-c", args["command"]], cwd=WORKDIR,
                            stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
                            encoding="utf-8", errors="replace",
                            start_new_session=True)     # own process group: a timeout kills it all
    try:
        out, _ = proc.communicate(timeout=TIMEOUT_S)
        status = f"exit code {proc.returncode}"
    except subprocess.TimeoutExpired:
        with contextlib.suppress(ProcessLookupError):
            os.killpg(proc.pid, signal.SIGKILL)
        out, _ = proc.communicate()
        status = (f"killed after {TIMEOUT_S} s. For a long job, run it in the background "
                  "with its output in a log file, then read the log in later calls")
    if not out.strip():
        out = "(the command printed nothing)"
    elif len(out) > KEEP:
        cut = len(out) - KEEP
        out = f"{out[:KEEP // 2]}\n[... {cut} characters cut from the middle ...]\n{out[-KEEP // 2:]}"
    return f"{status}\n{out}"


BASH = Tool(
    name="bash",
    description=(f"Run one bash command in the repository; returns the exit code and output. "
                 f"Commands are killed after {TIMEOUT_S} s. Output over {KEEP} characters "
                 "keeps its start and end."),
    input_schema={"type": "object",
                  "properties": {"command": {"type": "string", "description": "The command to run."}},
                  "required": ["command"]},
    run=bash,
)


class MalformedLimit(Middleware):
    """End the run after limit tool calls in a row that the harness could not run."""

    def __init__(self, limit: int = 3):
        self.limit, self.streak = limit, 0

    def after_tool(self, state: State, name: str, args: dict, result: str) -> str:
        self.streak = self.streak + 1 if result.startswith("error:") else 0
        if self.streak >= self.limit:
            state.stop_reason = "error"
        return result


SYSTEM = ("You are a coding agent in a git repository. Use the bash tool to read code, edit files "
          "and run checks. Run the tests before you say you are done. When the task is complete, "
          "reply with a two-line summary and no tool call.")

if __name__ == "__main__":
    task = " ".join(sys.argv[1:]) or "Run pytest -q, fix the first failing test, and rerun pytest -q."
    state = run(task, [BASH], system=SYSTEM, max_turns=30, max_usd=1.00,
                middleware=[MalformedLimit()], log_path="runs/session.jsonl")
    print(f"{state.stop_reason}: {state.turn} turns, ${state.cost_usd:.4f}, {state.usage}")
The bash tool runs as you

The commands the model writes run with your user's permissions and network access, in HX_WORKDIR. Use a throwaway clone inside a container or a VM, and an API key with a spending limit; Chapter 7 covers sandboxes and the network boundary.

Put the three files in place, point the script at a disposable repository with a failing test, and give it a dollar cap you are willing to lose.

Run it
bash
# Layout: hx/__init__.py (empty), hx/loop.py, hx/events.py, example_run.py
python3 -m venv .venv && source .venv/bin/activate
pip install anthropic
export ANTHROPIC_API_KEY=sk-ant-...        # your key

# A throwaway clone of a small repository with a failing test
git clone --depth 1 https://github.com/<you>/<small-repo>.git ~/scratch/target

HX_WORKDIR=~/scratch/target python example_run.py \
  "Run pytest -q, fix the first failing test, and rerun pytest -q."

python -m hx.events runs/session.jsonl   # rebuild every request from the log and check it

A run prints its stop reason, turns, cost and token totals and leaves the session in runs/session.jsonl. The output below is illustrative, written for this page rather than recorded, and its repository is made up.

A run, illustrative output
text
$ HX_WORKDIR=~/scratch/target python example_run.py "Run pytest -q, fix the first failing test, and rerun pytest -q."
done: 6 turns, $0.0453, {'input': 16712, 'output': 1187, 'cache_read': 0, 'cache_write': 0}

$ jq -c '[.turn, .kind]' runs/session.jsonl | head -n 9
[0,"start"]
[0,"message"]
[1,"request"]
[1,"response"]
[1,"message"]
[1,"tool"]
[1,"message"]
[2,"request"]
[2,"response"]

$ jq -c 'select(.kind == "tool") | [.turn, .data.args.command, (.data.result | split("\n")[0])]' runs/session.jsonl
[1,"pytest -q","exit code 1"]
[2,"sed -n '1,60p' dateparse/core.py","exit code 0"]
[3,"grep -rn \"def parse_date\" dateparse tests","exit code 0"]
[4,"python - <<'EOF2'\n...\nEOF2","exit code 0"]
[5,"pytest -q","exit code 0"]

$ python -m hx.events runs/session.jsonl
turn   1  ok    1 messages
turn   2  ok    3 messages
turn   3  ok    5 messages
turn   4  ok    7 messages
turn   5  ok    9 messages
turn   6  ok   11 messages

The event kinds trace the loop: request, response, assistant message, tool event, tool-result message, then the next request. Six turns made six requests, each two messages longer than the last, and the final command before the reply was the test suite exiting with code 0.

Log everything the model sees

DeepSeek Harness states the rule hx/events.py implements, “Model-visible means logged”: every model request must be reconstructable from the session log (architecture doc). hx also follows its second rule, “There is no privileged core to patch”, which is why later chapters plug in through hooks.

The log stores the parts of each request: start sets the model, system prompt and tools, each message appends, a rewrite replaces the history, and each request marks a send with its fingerprint. Storing parts keeps the log linear while the requests grow quadratically. requests() folds the parts back together, and verify rebuilds every request, checks its fingerprint and flags any turn whose predecessor is no longer its prefix, which is what costs a cache miss.

The figure steps through a session that crashes halfway, following the hx log; the labels under harness B are what Anthropic's Managed Agents calls the same two operations.

It opens on the final step; press NEXT to walk from the first event. Requests are folds of earlier events (thick boxes, arrows), harness A dies with request 2 in flight, and harness B rebuilds request 2, checks the fingerprint and resends it, without editing a line. The lost response is sampled again and can differ.

A log that can rebuild every request has four uses.

  • Debugging what the model saw. Anthropic's April 2026 postmortem describes idle-session thinking clearing (clear_thinking_20251015 with keep:1) that, through a bug, cleared thinking on every turn. It passed review, tests, automated verification and dogfooding, and took “over a week” to find. It changed what the model saw, which is what a request log records.
  • Finding cache breaks. Claude Code v2.1.275 fixed a restored memory file's “age note” that changed between requests and caused cache misses (changelog). The prefix check in verify flags that kind of bug on the first turn it happens.
  • Surviving crashes. Managed Agents keeps the session as an “append-only log” separate from the harness and reboots a crashed harness from it, which turned harnesses and containers from “pets” into “cattle” (Scaling Managed Agents). OpenHands reports that event-sourced state cut system-attributable failures by 61%, from 78.0 to 30.0 per 1,000 conversations in a 15-day rollout, with crash recovery under 20 ms (arXiv 2511.03690). Chapter 9 builds a restartable driver on the hx log.
  • Evaluating runs. The log is the transcript Chapter 12 grades and the raw material for turning a failure into a regression case (trace promotion and pass^k gates).

Log at the boundary the model sees: hx records tool results after after_tool has shaped them. For raw output, write a file and log its path, as Claude Code does with tool results over 50,000 characters (context window docs).

Observations the model can use

An observation is everything the model learns about what its call did; it cannot see the terminal. The harnesses that do well rewrite raw results into observations with a few consistent rules.

Raw resultWhat the model should readWho does this
The command printed nothinga line saying it ran and printed nothingSWE-agent replaces empty output with “Your command ran successfully and did not produce any output”
A nonzero exit statusthe exit code, every timeClaude Code treats exit code 1 as benign only for grep, rg, find, diff, test, git diff, git grep and a few others
Output too long for the windowthe head and the tail, with a note saying how much was cutCodex keeps head and tail around a “…N tokens truncated…” marker under a “Warning: truncated output” header, 10,000 tokens by default; Claude Code keeps about 30,000 characters inline, then gives a file path and a 2,000-character preview
A command that hangsa kill after a timeout, and a message saying what happened and what to tryClaude Code BASH_DEFAULT_TIMEOUT_MS 120000 (maximum 600000); mini-SWE-agent v2 defaults to 30 seconds
Output the agent may need latera file it can read or tail when it needs itCursor writes long shell and MCP output to files; Claude Code persists tool results over 50,000 characters
A failed checkthe error, plus the context needed to fix itSWE-agent's lint gate shows the error, the rejected edit and the original snippet

In the SWE-agent paper's ablations, leaving the original snippet out of the lint gate's message made the agent repeat the same command more often (arXiv 2405.15793). The Codex rule is in its output-truncation source, and Claude Code's limits are in its tools reference, where an oversized read now returns a PARTIAL view notice explaining offset and limit instead of an error. Cursor moved long outputs into files because truncation “can lead to data loss”, and reports fewer unnecessary summarizations (Dynamic context discovery). The timeouts are from Claude Code's environment variables.

Models read observations as instructions, so the wording shapes behavior. Under mini-swe-agent's 30-second timeout, GLM-5.3 and Kimi K3 learned to start long builds in the background and poll them with sleep 29; cat /tmp/build.log, as Scale and Reflection observed. The timeout message in example_run.py spells the same technique out, so a model that has not learned it gets it from the observation.

The bash tool in section 6 applies four of the six rules: it reports the exit code, names empty output, cuts the middle out of long output with a note, and kills a hung command with an explanation. Chapter 5 replaces it with tools that budget output in tokens, write long results to files and phrase errors as instructions.

The same loop, hosted

You do not have to own the loop. Anthropic's Agent SDK overview compares four routes: the Agent SDK (the loop in your process), the CLI, the Client SDK (your own tool loop, as in hx, or the beta tool runner) and Managed Agents (a hosted harness). OpenAI's range runs from the Codex SDK to a hosted Agents API.

The Claude Agent SDK is “a library that runs the Claude Code binary”, and TypeScript SDK v0.3.291 bundles Claude Code v2.1.291 (TypeScript reference). Two defaults matter when you compare it with hx. Since v0.1.0 the default system prompt “covers tool calling but omits the rest of the claude_code preset's content, including its security and safety instructions” (migration guide), and settings files load unless you pass an empty settingSources. The tabs run the example task through the SDK and the two CLIs.

The example task through the Agent SDK and the CLIs
loop_sdk.ts
typescript
// loop_sdk.ts: the example_run.py task through the Claude Agent SDK (Chapter 2).
// npm install @anthropic-ai/claude-agent-sdk, then: HX_WORKDIR=~/scratch/target npx tsx loop_sdk.ts
import { query } from "@anthropic-ai/claude-agent-sdk";

const SYSTEM =
  "You are a coding agent in a git repository. Use the bash tool to read code, edit files " +
  "and run checks. Run the tests before you say you are done. When the task is complete, " +
  "reply with a two-line summary and no tool call.";

async function main(task: string) {
  for await (const message of query({
    prompt: task,
    options: {
      model: "claude-sonnet-5-5",
      cwd: process.env.HX_WORKDIR ?? process.cwd(),
      systemPrompt: SYSTEM,
      tools: ["Bash"], // the only tool that exists in this session
      allowedTools: ["Bash"], // and it runs without a permission prompt
      permissionMode: "dontAsk", // anything not pre-approved is denied, never asked
      settingSources: [], // ignore user, project and local settings files
      maxTurns: 30,
      maxBudgetUsd: 1.0,
    },
  })) {
    if (message.type === "assistant") {
      for (const block of message.message.content) {
        if (block.type === "tool_use") console.log("bash:", JSON.stringify(block.input));
      }
    } else if (message.type === "result") {
      // subtype: success | error_max_turns | error_max_budget_usd | error_during_execution
      console.log(message.subtype, message.num_turns, "turns", "$" + message.total_cost_usd.toFixed(4));
    }
  }
}

main(process.argv.slice(2).join(" ") || "Run pytest -q, fix the first failing test, and rerun pytest -q.");

tools decides which tools exist, while allowedTools only pre-approves them; the reference says it “does not restrict Claude to only these tools.” A permissionMode of dontAsk denies anything not pre-approved instead of prompting, which an unattended run needs, and a capped run ends with a result whose subtype names the cap. The SDK brings Claude Code's tools, compaction, hooks and session files, and in exchange part of each request comes from the bundled binary instead of your code.

Anthropic's Managed Agents moves the loop off your machine and splits it three ways: the session (an “append-only log”), the harness (the loop that calls Claude and routes tool calls) and the sandbox (where tools run), joined by small interfaces such as execute(name, input) returning a string, wake(sessionId) and getEvents(). Anthropic calls it “a meta-harness” and notes that “The harness doesn't know whether the sandbox is a container, a phone, or a Pokémon emulator.”

The split came from a first design that ran all three in one container, a pet whose only debugging window was the WebSocket event stream. Now a dead container surfaces as a tool-call error passed back to Claude, a crashed harness reboots from the session log, and “The harness is never made aware of any credentials.” Provisioning containers only when a tool call needs one cut p50 time to first token by roughly 60%, per Anthropic. It is in public beta at token rates plus $0.08 per session-hour of active runtime (launch post), behind the managed-agents-2026-04-01 beta header, and not eligible for Zero Data Retention or a HIPAA BAA (docs).

OpenAI's Agents API (public beta since 2026-09-10) offers “the same harness and infrastructure that powers Codex”. A session starts with client.beta.agents.sessions.create, which takes an agent (model, tools, an optional multi_agent block), vault_ids and an environment of type openai_hosted or self_hosted; compaction, tool search and programmatic tool calling are managed, with “no additional fees” beyond tokens and tools. The launch post's customer numbers are claims, such as Hypha's 86% fewer failed responses “by separating the agent harness from the sandbox”. The Codex SDK drives the same harness from your own code (new Codex(), startThread(), thread.run()).

Claude Agent SDKAnthropic Managed AgentsOpenAI Agents API
Where the loop runsyour process, through the bundled Claude Code binaryAnthropic's hosted harnessOpenAI's hosted Codex harness
Where tools runyour machine, with a sandbox optionan Environment: Anthropic's cloud or a self-hosted sandboxan openai_hosted or self_hosted environment
Session statesession files on disk (persistSession); sessionStore mirrors them elsewherean append-only session log; wake(sessionId), getEvents()hosted sessions (client.beta.agents.sessions)
Context handlingClaude Code's compactioncompaction, reported as agent.thread_context_compacted eventsmanaged automatic compaction
Price beyond tokensnone; it is a library$0.08 per session-hour of active runtime“no additional fees” beyond tokens and tools
Status, October 2026TypeScript v0.3.291, Python v0.2.163public beta since 2026-04-08; not eligible for ZDR or a HIPAA BAApublic beta since 2026-09-10

Whichever you pick, ask it the three questions hx answers in code: where the session log lives, where tools run, and which caps end a run.

Where the minimal loop breaks

The minimal loop works when the model is strong, the task fits one context window and nothing the model might do is dangerous. Real work breaks those conditions in five recurring ways, each with a chapter of its own.

Five ways the minimal loop breaks

Step 1 / 5
1. Context overflow and rot (Chapters 3 and 4)
Every turn appends the response and the full observation, and nothing ever leaves. In arXiv 2609.20804, runs without context management overflowed the window on 78.7% of SWE-bench tasks at 32K tokens and 8.7% at 128K, and managed runs never did. When the input no longer fits, the API returns a 400 invalid_request_error (“prompt is too long”) and hx stops with error. In arXiv 2608.26218, shortening old tool results and handling stalls in one harness raised mean fail-to-pass from 28% to 49% on 169 SWE-bench Verified tasks at a 20,480-token window. Chapters 3 and 4 budget the window and shrink what fills it.

None of these fixes needs a bigger core. Each later chapter adds one hx module through the four hooks, and Chapter 13 assembles them into a background agent that goes from an issue to a reviewed pull request.

Chapter 2 in one page

Key takeaways
8 items
  • 1A harness decides what the model sees on every request and what happens to everything it asks for.
  • 2Codex assembles instructions, tools, permissions, the AGENTS.md chain (up to 32 KiB), environment and then the user message, with stable layers first because caches hit only on exact prefixes.
  • 3Every request resends the transcript: 20K to 120K tokens over 40 turns is about 2.8M input tokens, about 96% of them cacheable with a perfect prefix. Append changes instead of editing the prefix.
  • 4Minimal loops are strong baselines: mini-SWE-agent with one bash tool leads Terminal-Bench 3.0 with Opus 5 at 42.7%, and Pi's prompt and four tools fit in under 1,000 tokens.
  • 5Keep a turn cap and a dollar cap priced from usage in the core; completion checks, repeat detectors and malformed-call limits are middleware whose firing rates are health metrics.
  • 6Log the request as the model saw it, rebuild every request from the log, and check fingerprints and prefixes.
  • 7Observations carry the exit code, never come back empty, say what was cut, and tell the model what to do after a timeout.
  • 8The minimal loop breaks five ways: overflow (Chapters 3 and 4), output flooding (5), false completion (6), destructive actions (7) and lost state across windows (9).

What to do on Monday: copy hx/loop.py, hx/events.py and example_run.py, point the script at a throwaway clone of a repository with one failing test, and set max_usd to 1.00. Run it, run python -m hx.events runs/session.jsonl, and read the requests the model saw, turn by turn. Then check whether the last command before the final reply was a passing test run.

Chapters 3 to 13: Building, Securing, Scaling and Measuring the Harness

Unlock context budgeting, caching and compaction, the action space, feedback gates and evaluators, permissions and sandboxes, orchestration and long-running state, the repository as the harness, review and merge at agent speed, the eval rig for harness changes, a background coding agent built end to end, and the landscape and sources appendix.

Join The Agent Foundry to unlock chapters 3 to 13 (context and caching, tools, feedback and evaluators, permissions and sandboxes, orchestration, long-running state, the repository as the harness, review and merge, measuring a harness, and a background coding agent built end to end), the landscape and sources appendix, and every future book on release.

Enter your email to continue. We'll send a one-click sign-in link and bring you back here.