Paper Reading / Harness Engineering

论文精读:十一个编码 harness 的源码级解剖——七个子系统、13 条跨系统观察、29 个复用模式,以及两条经受住复核的缺席结论:没有通用 agent 框架,也没有向量检索查代码。

harness 解剖学

把十一个生产级编码 harness 的源码拆开,看它们到底由什么组成。

这篇论文做了一件此前没人做过的事:把十一个生产级编码 harness 的源码全部拆开,用同一张图纸摆好零件,然后告诉你哪些零件所有人都装了、哪些只有一家装、哪些看起来重要其实没人装。它的两条结论最反直觉——循环的复杂度不预测能力,而约四百万行代码里没有任何系统使用通用 agent 框架或向量检索查代码。

The paper does something nobody had done: it opens up eleven production coding harnesses at the source level and lays the parts out on one diagram, so you can see which parts everyone installs, which only one vendor installs, and which look important but nobody uses. Two findings are counter-intuitive — loop sophistication does not predict capability, and across roughly four million lines no system uses a general-purpose agentic framework or vector retrieval over code.

阅读结论 Verdict 设计自查表 / 不是性能依据 A design checklist, not a performance claim 七子系统与 18 条建议可直接对照自己的代码逐条自查;但全文只看源码、不做运行时测量,也没有同任务横评。 The seven subsystems and 18 recommendations map cleanly onto your own code; but the study reads source rather than measuring runtimes, and never runs the eleven systems on one task set.
11 + 1 十一个编码 harness(四家厂商自有与七个开源)加一个元 harness 对照点 Eleven coding harnesses (four vendor-native, seven open source) plus one meta-harness as a contrast point
7 / 13 / 29 / 18 七个子系统、13 条跨系统观察、29 个复用模式、18 条设计建议 Seven subsystems, 13 cross-cutting observations, 29 recurring patterns, 18 design recommendations
0 / 11 使用通用 agent 框架的系统数,以及用向量嵌入检索代码的系统数 Systems using a general-purpose agentic framework, and systems using vector embeddings to retrieve code
论文 Paper

Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents

Paul Barbaste(负责人与通讯作者)、Tristan Darrigol、Germain Vu、Tom Wiltberger;Inclusive Brains 与 Wavestone AI Lab。

Paul Barbaste (lead and corresponding author), Tristan Darrigol, Germain Vu and Tom Wiltberger, from Inclusive Brains and Wavestone AI Lab.

来源 Source

arXiv:2609.00006v1

2026-07-15 提交,83 页、7 图、18 表;2026 年 4 月版(8 个系统)的实质性扩展版。

Submitted 2026-07-15, 83 pages, 7 figures, 18 tables; a substantially expanded second edition of an April 2026 study over eight systems.

官方资源 Artifacts

论文未给出代码或数据

No code or data link provided

但每个系统都固定到 release tag 与 commit,并保留 4 月快照,读者可自行复核源码。

Every system is pinned to a release tag and commit, and the April snapshots are retained, so the source claims can be re-checked independently.

问题:agent 是模型加 harness

Problem: an agent is a model plus a harness

论文的出发点是那句代数式:Agent = Model + Harness。模型提供智能,harness 把这份智能转成工作——通过循环、工具、上下文管理、安全控制、编排与扩展面。「harness engineering」这个词在 2026 年 2 月进入流通,五个月内就长成了有实践指南、形式定义与自动演化系统的学科。

The paper starts from the algebra: Agent = Model + Harness. The model supplies the intelligence; the harness turns it into work through a loop, tools, context management, safety controls, orchestration and extension surfaces. The term harness engineering entered circulation in February 2026 and became a discipline with practitioner guides, formal definitions and automated-evolution systems within five months.

缺的是一份参照物:基于生产源码说明 harness 到底是什么、由什么组成、领先实现差在哪里。既有文献都没做到——benchmark 榜单只报分数不碰架构,概念性分类只讲模式没有实现,已有的源码级分类又排除了定义前沿的厂商原生系统。

What was missing is a reference: an account grounded in production source code of what a harness is, what it is made of, and where leading implementations differ. Existing literature reaches neither — benchmark surveys report scores without touching architecture, conceptual taxonomies describe patterns without implementations, and the one concurrent source-level taxonomy excludes the vendor-native systems that define the frontier.

论文给了四个边界情形来固定研究对象:harness 不是 scaffold(scaffold 指结构性代码,harness 指交付的运行时产物)、不是 agentic framework(框架是你导入的库,harness 是你在里面工作的运行时)、不是 evaluation harness(那个的方向是包住 agent 去跑任务)、不是 orchestrator(编排器协调别人,自己并不实现编辑循环)。

Four boundary cases fix the object of study. A harness is not a scaffold (scaffold names the structural code, harness the shipped runtime artifact), not an agentic framework (a framework is a library you import, a harness is a runtime you work inside), not an evaluation harness (that wraps an agent to run tasks, the opposite direction), and not an orchestrator (which coordinates others and implements no editing loop).

七个子系统与实现区间

Seven subsystems and their range

论文主张:从约百行的研究基线到百万行的生产 CLI,每个系统都必须对这七件事表态——哪怕表态是刻意缺席(Aider 的编排就是如此)。

The claim is that from a 100-line research baseline to a million-line production CLI, every system must take a position on the same seven things — even when the position is deliberate absence, as with Aider's orchestration.

harness 的七个子系统:agent 循环居中,LLM 集成、记忆与上下文、工具与动作、安全与权限、编排、扩展性环绕,另有接口层与会话基底两个横切面 The seven subsystems of a harness: the agent loop at the centre, with LLM integration, memory and context, tools and actions, safety and permissions, orchestration and extensibility around it, plus interface and session surfaces
图 1 · 七个子系统,外加两个横切面:接口层(TUI / CLI / IDE / SDK / HTTP 服务)与会话基底(transcript、持久化、resume 与 fork)。依据原文 Sec.2.3、Figure 1 与 Table 1 重绘。
Figure 1 - Seven subsystems plus two cross-cutting surfaces: the interface layer (TUI, CLI, IDE, SDK, HTTP server) and the session substrate (transcripts, persistence, resume and fork). Redrawn from Sec.2.3, Figure 1 and Table 1.
七个子系统各自的最小实现与最大实现对照 Minimal and maximal observed implementations for each of the seven subsystems
图 2 · 同一格子上的两端差三个数量级。最小实现几乎全部落在 Mini-SWE-Agent(约 100 行),最大实现分散在不同系统——说明不同系统在同一个决策面上选了不同的投资点。数据来源:原文 Table 1。
Figure 2 - Three orders of magnitude between the two ends of the same cell. Minimal implementations land almost entirely on Mini-SWE-Agent (~100 lines) while maximals spread across systems, showing different systems invest in different places along the same decision surface. Source: Table 1.
子系统 Subsystem 职责 Role 最大实现 Maximal form
Agent 循环 Agent loop 交替推理与执行,拥有停止条件与失败恢复 Alternates inference with action; owns stop conditions and recovery OpenHands:事件溯源对话,持久事件日志与并行动作批次 OpenHands: event-sourced conversation over a persistent log, with parallel action batches
LLM 集成 LLM integration 讲 provider 协议,组装提示,管理缓存与路由 Speaks provider protocols; assembles prompts; caching and routing Hermes:五套自有传输与 29 个 provider 档案;Codex:服务端下发模型目录 Hermes: five owned transports and 29 provider profiles; Codex: server-delivered model catalog
工具与动作 Tools and actions 定义并执行 agent 能做的事,首要是文件编辑 Defines and executes what the agent can do, file editing above all Claude Code:43 个类型化工具与延迟加载;Codex:工具调用作为 V8 执行代码 Claude Code: 43 typed tools with deferred loading; Codex: tool calls as V8-executed code
记忆与上下文 Memory and context 配给上下文窗口,跨回合与会话持久知识 Rations the context window; persists knowledge across sessions Codex:agent 维护的跨会话记忆管线;Gemini CLI:图式上下文蒸馏 Codex: agent-maintained cross-session memory pipeline; Gemini CLI: graph-based context distillation
安全与权限 Safety and permissions 决定什么能跑、什么要问、什么禁止,并隔离执行 Decides what runs, what asks, what is forbidden; isolates execution Codex:策略规则与审批评审,加三平台 OS 沙箱 Codex: policy rules, an LLM approval reviewer, and a three-platform OS sandbox
编排 Orchestration 派生并协调子 agent,连接其他 agent Spawns and coordinates sub-agents; connects to other agents Claude Code:递归组合;Omnigent:跨厂商协调(元层) Claude Code: recursive composition; Omnigent: cross-vendor coordination at the meta layer
扩展性 Extensibility 让用户与生态增加能力:配置、hooks、技能、插件、MCP Lets users and ecosystems add capability: config, hooks, skills, plugins, MCP Pi:一切皆扩展的运行时;Codex:市场分发的插件 Pi: an everything-is-an-extension runtime; Codex: marketplace-distributed plugins

十三条跨系统观察

Thirteen cross-cutting observations

观察的排序大致按子系统的顺序,每条都指向具体系统与源码位置。下面挑出对工程决策影响最大的几条。

The observations roughly follow the subsystem order, and each points at named systems and source locations. Here are the ones that most affect engineering decisions.

# 观察 Observation 关键证据 Key evidence
1 循环复杂度不预测能力;代码量主要花在循环之外 Loop sophistication does not predict capability; the mass sits outside the loop Mini-SWE-Agent 约百行循环与 OpenHands 事件溯源引擎报告结果同区间;OpenCode 非测试源码约五分之三是客户端 Mini-SWE-Agent's ~100-line loop reports results in the same range as OpenHands' event-sourced engine; roughly three-fifths of OpenCode's non-test source is clients
2 厂商原生优化的门槛是谁承担逐 provider 的条件代码成本 Provider-native optimisations hinge on who pays the per-provider conditional-code cost Hermes 手写五套传输与 29 个 provider 档案;Pi 把每轮缓存未命中的美元浪费当一等指标;OpenCode 同时发出六家缓存方言 Hermes hand-rolls five transports and 29 provider profiles; Pi treats per-turn cache-miss dollar waste as a first-class metric; OpenCode emits six cache dialects at once
3 提示修辞在经验收敛处收敛,随后随信任校准变薄;十一系统都没有政策级拒绝语言 Prompt rhetoric converges where experience converges and thins as trust calibrates; no policy-level refusal language anywhere 「不要镀金」在六个独立 scaffold 里措辞近乎同构;4 月的禁止自动提交规则到 7 月被反转或删除 Anti-gold-plating language appears near-isomorphically in six independently developed scaffolds; the April no-commit rule was reversed or dropped by July
4 文件编辑策略是代码修改准确率的最重要决定因素之一,且按模型多态 File-editing strategy is a top determinant of accuracy, and is model-aware Aider 的 prompt-class 工厂;OpenCode 按模型换工具集;Mistral Vibe 一个季度内从模糊匹配转向精确匹配 Aider's prompt-class factory; OpenCode swapping toolsets per model; Mistral Vibe moving from fuzzy to exact matching within one quarter
5 持久记忆取代上下文压缩成为前沿;四类记忆写入路径并存 Persistent memory replaced compaction as the frontier; four write-path governance models coexist 压缩已收敛(7/11 阈值触发 LLM 摘要);路径分为 agent 自主维护、人工把关、模型直写但有界、回合前召回 Compaction has converged (7/11 use threshold-triggered summarisation); paths split into agent-maintained, human-gated, bounded-direct and pre-turn recall
6 OS 级沙箱昂贵,而且是选择而非规模的必然结果 OS sandboxing is expensive, and a choice rather than a consequence of scale 4 月的「规模蕴含沙箱」相关性被打破:Hermes 与 OpenCode 属最大系统却零 OS 级隔离 April's size-implies-sandbox correlation broke: Hermes and OpenCode are among the largest systems with zero OS-level isolation
7 coordinator-worker 形态独立涌现;子 agent 协调大多仍在进程内 Coordinator-worker emerged independently; sub-agent coordination mostly stays in-process 四个厂商系统全部出现该形态;九分之八的多 agent 系统其子 agent 协调用进程内原语 All four vendor systems show the shape; eight of nine multi-agent systems coordinate sub-agents with in-process primitives
8 Skills 超过 MCP 成为采用最广的扩展标准 Skills overtook MCP as the most-adopted extensibility standard SKILL.md 9/11 对 MCP 8/11;延迟加载 8/9;出现注册表、信任层级、隔离与 agent 自撰技能 SKILL.md 9/11 against MCP 8/11; deferred loading 8/9; registries, trust tiers, quarantine and the first agent-authored skills
11 协议位置从两角色变三角色:ACP 新增 harness 托管 Protocol placement went from two roles to three: ACP added harness hosting ACP 6/11;OpenHands 把 Claude Code、Codex、Gemini CLI 当可互换后端;A2A 仍只有 Gemini CLI ACP ships in 6/11; OpenHands runs Claude Code, Codex or Gemini CLI as interchangeable backends; A2A remains Gemini CLI's bet alone
12 平台化转变已完成:竞争单位从循环移到循环之外的生态面 The platform turn is complete: the competitive unit moved from the loop to the ecosystem around it SDK 化发布、框架厂商发 harness、市场与信任层级、跨厂商会话导入、MDM 治理、元 harness 编排十一家厂商 SDK-shaped releases, framework vendors shipping harnesses, marketplaces and trust tiers, cross-vendor session importers, MDM governance, and a meta-harness orchestrating eleven vendors

两条经受过复核的缺席

Two absences that survived re-auditing

这两条是全文最可执行的结论,因为它们说的是不要做什么。作者检查了每个依赖清单,并在三种语言里 grep 了十余个通用框架的 import——跨十二棵树复核两次,相隔三个月;语料还经历了一次三倍扩展。结论没有被撼动。

These are the most actionable findings because they are about what not to build. Every dependency manifest was inspected and every tree grepped for imports of a dozen general-purpose frameworks — across twelve trees, twice, three months apart, with the corpus tripling in between. The result held.

两条缺席:没有系统使用通用 agent 框架,也没有系统用向量嵌入检索代码;右侧为采用率与替代方案 Two absences: no system uses a general-purpose agentic framework, and none uses vector embeddings to retrieve code; adoption rates and substitutes on the right
图 3 · 两条缺席与它们实际采用的技术栈。检查覆盖 LangChain、LangGraph、AutoGen、CrewAI、Pydantic AI、Genkit、Semantic Kernel、Google ADK 等;检索一律是 ripgrep、tree-sitter、glob 与自动发现的 Markdown 上下文。数据来源:原文 Observation 8、9(13.2)与 11。
Figure 3 - The two absences and the stack actually in use. The sweep covered LangChain, LangGraph, AutoGen, CrewAI, Pydantic AI, Genkit, Semantic Kernel, Google ADK and more; retrieval is consistently ripgrep, tree-sitter, glob and auto-discovered Markdown context. Source: Observations 8, 9 (13.2) and 11.

作者给出的解释是失败模式的权重:生产 harness 会修改真实代码,因此可调试性与提示透明度压过框架复用带来的便利。他们也把这条结论明确标为「结构上保守」——没有追内部 fork、动态 importlib 与 require 加载的插件、以及转译产物,所以不能说这种用法不存在,只能说没被这轮审计发现。

The explanation given is the weight of the failure mode: production harnesses mutate real code, so debuggability and prompt transparency outweigh the convenience of framework reuse. The authors also label the finding structurally conservative — internal forks, dynamically imported plugins and transpiled distributions were not traced, so the claim is that such use was not found, not that it does not exist.

18 条设计建议与 90 行骨架

18 recommendations and a 90-line scaffold

论文的最后一节是给从业者的清单:每条建议都引用支持它的观察、点名实现它的系统,并指出可能让人做出不同选择的取舍。

The last section is a practitioner's guide: each recommendation cites the observations that support it, names the systems that implement it, and identifies the trade-off that might justify a different choice.

18 条设计建议按子系统分组,红框为与行业常态相反的红线式建议 The 18 design recommendations grouped by subsystem, with red-line recommendations outlined
图 4 · 18 条建议按子系统分组。红框三条是「与常见做法相反」的红线:不要为代码建 RAG、不要在运行时用通用 agent 框架、不要为代码建向量检索层。数据来源:原文 Sec.16。
Figure 4 - The 18 recommendations grouped by subsystem. The outlined three run against common practice: no RAG over code, no general-purpose framework in the runtime, and no vector-embedding retrieval layer for code. Source: Sec.16.

论文还给了一个约 90 行的最小可用 harness,直接实现 18 条中的 10 条、其余兼容。它没有框架依赖、没有 RAG、没有向量库、没有多 agent 编排、也没有沙箱——正好与两条缺席结论一致。作者明确说这是推测(无证明):这个骨架在前沿模型上会接近 Mini-SWE-Agent 的数字,而再往上走主要是模型能力问题,不是 scaffold 问题。

The paper also ships a roughly 90-line minimum-viable harness that realises ten of the eighteen recommendations directly and stays compatible with the rest. It has no framework dependency, no RAG, no vector store, no multi-agent orchestration and no sandbox — exactly in line with the twin absences. The authors are explicit that the follow-on claim is a conjecture, not a proof: that this scaffold would approach Mini-SWE-Agent's numbers on a frontier model, and that moving beyond those numbers is mostly a model-capability question rather than a scaffold question.

90 天纵向与平台化

Ninety days of evolution, and the platform turn

因为 4 月版的八个系统是重新钉版本而非被替换,论文手里有一份受控的纵向样本:同一批 harness 的源码 diff 跨越一个季度。结论是「收敛变成了模仿」——Codex 逐字采用 Claude Code 的 hook 词汇并发布会话与设置导入器;OpenHands 读取 Claude Code 的插件格式;行为政策从提示散文迁移到配置(feature flag 与服务端 A/B 测试)。4 月版的三条观察被新证据实质修订。

Because the eight systems from the April edition were re-pinned rather than replaced, the study holds a controlled longitudinal sample: the same harnesses, source-diffed across one quarter. The result is convergence becoming imitation — Codex adopts Claude Code's hook vocabulary verbatim and ships session and settings importers, OpenHands reads Claude Code's plugin format, and behavioural policy migrates from prompt prose into configuration and server-side A/B tests. Three of the April observations were substantively revised on new evidence.

由此得出的判断是:编码 harness 在 2026 年上半年完成了从工具到平台的转变。证据是具名产物:harness 以可导入 SDK 的形式发布,而框架厂商反过来发布自己的 harness;能力分发有了市场、注册表与信任层级;厂商为彼此的磁盘状态写导入器;出现了 MDM 级别的治理层;agent 本身以 OpenAI 兼容端点暴露;一个元 harness 在一个 API 后面编排十一家厂商的 harness,重新实现其中昂贵的部分、套利专有的部分。作者因此把竞争单位从 agent 循环重新定位到循环之外的生态面。

The resulting judgement is that the coding harness completed its turn from tool to platform in the first half of 2026. The evidence is a list of named artifacts: harnesses ship as importable SDKs while framework vendors ship harnesses; capability distribution gained marketplaces, registries and trust tiers; vendors write importers for each other's on-disk state; MDM-grade governance appeared; the agent itself became addressable behind an OpenAI-compatible endpoint; and one meta-harness orchestrates eleven vendor harnesses behind a single API, re-implementing the expensive parts and arbitraging the proprietary ones. The authors therefore relocate the competitive unit from the agent loop to the ecosystem around it.

批判性评估

Critical assessment

强证据 Strong evidence 每条观察、模式与建议都锚定到具体模块与版本,作者刻意不做行号引用以避免快速过期,改为把 release tag 与 commit 固定下来,读者可自行复核。两条缺席结论经受住了三重检验:三倍语料扩展、三种语言的依赖清单与 import grep、以及相隔三个月的两次复核。纵向对比是同语料两次快照,而非不同系统的横向拼接,这让「模仿」这一结论有可控的比较基础。采用率是可数的(Skills 9/11、MCP 8/11、延迟加载 8/9),比纯定性描述更可检验。作者还主动披露了方法论局限,包括最弱的一支。 Every observation, pattern and recommendation is anchored to a named module and version. The authors deliberately avoid line-number citations because they decay, and instead pin release tags and commits so readers can re-check the claims. The twin absences survived three tests: a threefold corpus expansion, dependency manifests plus import greps across three languages, and two re-audits three months apart. The longitudinal comparison uses two snapshots of the same corpus rather than a stitched cross-section, which gives the imitation finding a controlled basis. Adoption rates are countable (skills 9/11, MCP 8/11, deferred loading 8/9), making them more checkable than prose. The authors also volunteer their methodological limits, including their weakest link.
中等或弱证据 Weaker evidence 全文只看源码,不做运行时测量,因此无法回答「哪套架构更快、更省、更可靠」;引用的 benchmark 数字全部来自各系统自己的文档,作者也提醒最后一位小数不是重点。架构评分包含主观判断,作者自己承认并欢迎异议;七子系统的归格(记忆与上下文合并为一格)会掩盖系统间的真实差异。两条缺席结论是结构上保守的,未追内部 fork、动态加载的插件与转译产物。没有同任务横向执行实验,跨系统比较无法转化为性能排序——作者把这一步列为未来工作并承认成本高一个数量级。Claude Code 一支依赖 2026 年 3 月公开流出的源码快照,与另外十个可从公开 Git 历史重建的系统存在不对称。此外全文在 Claude Code CLI 的实质性协助下完成,是「研究编码 harness 的论文由编码 harness 辅助产出」的案例,作者按会议政策做了披露。 The study reads source rather than measuring runtimes, so it cannot say which architecture is faster, cheaper or more reliable; every benchmark figure quoted comes from the systems' own documentation, and the authors note that the last decimal place is not the point. Architectural scoring involves judgement calls, which the authors acknowledge and invite disagreement on; collapsing memory and context into one cell also hides real differences between systems. The absences are structurally conservative, since internal forks, dynamically loaded plugins and transpiled distributions were not traced. There is no head-to-head execution study, so cross-system comparisons cannot become a performance ranking — the authors list it as future work and concede it costs an order of magnitude more effort. The Claude Code analysis rests on a publicly circulated March 2026 source snapshot, an asymmetry with the ten systems rebuildable from public Git history. Finally, the paper was written with substantial assistance from the Claude Code CLI, making it a case of a harness study produced with a harness, disclosed in line with conference policies.
结论是否超出证据 Does the claim exceed the evidence 「七子系统是规范解剖」作为组织框架成立且自洽,但它是分析工具而非自然规律。「循环复杂度不预测能力」方向成立,有 Mini-SWE-Agent 的存在性证明支撑;但它用的是各系统自报数字、模型与日期都不同,不能读成严格的等价性。「两条缺席」成立且保守,作者把检验范围与盲区都写清楚了。「平台化转变已完成」的证据是具名产物清单,这类清单声明会随时间过期,作者也把它标为「清单类」,方向上站得住但应带日期引用。「与 Anthropic 工程系列高度吻合」被作者自己标为因果未定。 The seven-subsystem anatomy holds as an organising frame, but it is an analytical tool rather than a law of nature. The claim that loop sophistication does not predict capability is directionally supported by the Mini-SWE-Agent existence proof, but it rests on self-reported numbers from different models and dates and cannot be read as strict equivalence. The twin absences hold and are conservative, with both the scope of the check and its blind spots stated. The claim that the platform turn is complete rests on a list of named artifacts — list-shaped claims decay, the authors label it as such, and it should be cited with a date. The alignment with Anthropic's engineering series is marked by the authors themselves as causally unestablished.

这篇论文的价值在于把「harness 到底是什么」从讨论变成了可核对的清单:七子系统给了设计自查的格子,29 个模式给了实现选项的目录,18 条建议给了起步顺序,两条缺席给了两个常被默认接受的做法以强反例。

The value here is turning "what is a harness" from a discussion into a checkable inventory: seven subsystems give the audit cells, 29 patterns give the option catalogue, 18 recommendations give the starting order, and two absences give strong counterexamples to two widely assumed practices.

一句话:循环的复杂度不预测能力,代码量主要花在循环之外;生产编码 harness 不用通用框架,也不用向量检索查代码——按这三条去检查自己的 harness,比按榜单排名去抄更有效。 In one line: loop sophistication does not predict capability, the mass sits outside the loop, and production coding harnesses use neither a general-purpose framework nor vector retrieval over code — auditing your own harness against those three beats copying a leaderboard.