harness 解剖学
把十一个生产级编码 harness 的源码拆开,看它们到底由什么组成。
这篇论文做了一件此前没人做过的事:把十一个生产级编码 harness 的源码全部拆开,用同一张图纸摆好零件,然后告诉你哪些零件所有人都装了、哪些只有一家装、哪些看起来重要其实没人装。它的两条结论最反直觉——循环的复杂度不预测能力,而约四百万行代码里没有任何系统使用通用 agent 框架或向量检索查代码。
The paper does something nobody had done: it opens up eleven production coding harnesses at the source level and lays the parts out on one diagram, so you can see which parts everyone installs, which only one vendor installs, and which look important but nobody uses. Two findings are counter-intuitive — loop sophistication does not predict capability, and across roughly four million lines no system uses a general-purpose agentic framework or vector retrieval over code.
问题:agent 是模型加 harness
Problem: an agent is a model plus a harness
论文的出发点是那句代数式:Agent = Model + Harness。模型提供智能,harness 把这份智能转成工作——通过循环、工具、上下文管理、安全控制、编排与扩展面。「harness engineering」这个词在 2026 年 2 月进入流通,五个月内就长成了有实践指南、形式定义与自动演化系统的学科。
The paper starts from the algebra: Agent = Model + Harness. The model supplies the intelligence; the harness turns it into work through a loop, tools, context management, safety controls, orchestration and extension surfaces. The term harness engineering entered circulation in February 2026 and became a discipline with practitioner guides, formal definitions and automated-evolution systems within five months.
缺的是一份参照物:基于生产源码说明 harness 到底是什么、由什么组成、领先实现差在哪里。既有文献都没做到——benchmark 榜单只报分数不碰架构,概念性分类只讲模式没有实现,已有的源码级分类又排除了定义前沿的厂商原生系统。
What was missing is a reference: an account grounded in production source code of what a harness is, what it is made of, and where leading implementations differ. Existing literature reaches neither — benchmark surveys report scores without touching architecture, conceptual taxonomies describe patterns without implementations, and the one concurrent source-level taxonomy excludes the vendor-native systems that define the frontier.
论文给了四个边界情形来固定研究对象:harness 不是 scaffold(scaffold 指结构性代码,harness 指交付的运行时产物)、不是 agentic framework(框架是你导入的库,harness 是你在里面工作的运行时)、不是 evaluation harness(那个的方向是包住 agent 去跑任务)、不是 orchestrator(编排器协调别人,自己并不实现编辑循环)。
Four boundary cases fix the object of study. A harness is not a scaffold (scaffold names the structural code, harness the shipped runtime artifact), not an agentic framework (a framework is a library you import, a harness is a runtime you work inside), not an evaluation harness (that wraps an agent to run tasks, the opposite direction), and not an orchestrator (which coordinates others and implements no editing loop).
七个子系统与实现区间
Seven subsystems and their range
论文主张:从约百行的研究基线到百万行的生产 CLI,每个系统都必须对这七件事表态——哪怕表态是刻意缺席(Aider 的编排就是如此)。
The claim is that from a 100-line research baseline to a million-line production CLI, every system must take a position on the same seven things — even when the position is deliberate absence, as with Aider's orchestration.
| 子系统 | Subsystem | 职责 | Role | 最大实现 | Maximal form |
|---|---|---|---|---|---|
| Agent 循环 | Agent loop | 交替推理与执行,拥有停止条件与失败恢复 | Alternates inference with action; owns stop conditions and recovery | OpenHands:事件溯源对话,持久事件日志与并行动作批次 | OpenHands: event-sourced conversation over a persistent log, with parallel action batches |
| LLM 集成 | LLM integration | 讲 provider 协议,组装提示,管理缓存与路由 | Speaks provider protocols; assembles prompts; caching and routing | Hermes:五套自有传输与 29 个 provider 档案;Codex:服务端下发模型目录 | Hermes: five owned transports and 29 provider profiles; Codex: server-delivered model catalog |
| 工具与动作 | Tools and actions | 定义并执行 agent 能做的事,首要是文件编辑 | Defines and executes what the agent can do, file editing above all | Claude Code:43 个类型化工具与延迟加载;Codex:工具调用作为 V8 执行代码 | Claude Code: 43 typed tools with deferred loading; Codex: tool calls as V8-executed code |
| 记忆与上下文 | Memory and context | 配给上下文窗口,跨回合与会话持久知识 | Rations the context window; persists knowledge across sessions | Codex:agent 维护的跨会话记忆管线;Gemini CLI:图式上下文蒸馏 | Codex: agent-maintained cross-session memory pipeline; Gemini CLI: graph-based context distillation |
| 安全与权限 | Safety and permissions | 决定什么能跑、什么要问、什么禁止,并隔离执行 | Decides what runs, what asks, what is forbidden; isolates execution | Codex:策略规则与审批评审,加三平台 OS 沙箱 | Codex: policy rules, an LLM approval reviewer, and a three-platform OS sandbox |
| 编排 | Orchestration | 派生并协调子 agent,连接其他 agent | Spawns and coordinates sub-agents; connects to other agents | Claude Code:递归组合;Omnigent:跨厂商协调(元层) | Claude Code: recursive composition; Omnigent: cross-vendor coordination at the meta layer |
| 扩展性 | Extensibility | 让用户与生态增加能力:配置、hooks、技能、插件、MCP | Lets users and ecosystems add capability: config, hooks, skills, plugins, MCP | Pi:一切皆扩展的运行时;Codex:市场分发的插件 | Pi: an everything-is-an-extension runtime; Codex: marketplace-distributed plugins |
十三条跨系统观察
Thirteen cross-cutting observations
观察的排序大致按子系统的顺序,每条都指向具体系统与源码位置。下面挑出对工程决策影响最大的几条。
The observations roughly follow the subsystem order, and each points at named systems and source locations. Here are the ones that most affect engineering decisions.
| # | 观察 | Observation | 关键证据 | Key evidence |
|---|---|---|---|---|
| 1 | 循环复杂度不预测能力;代码量主要花在循环之外 | Loop sophistication does not predict capability; the mass sits outside the loop | Mini-SWE-Agent 约百行循环与 OpenHands 事件溯源引擎报告结果同区间;OpenCode 非测试源码约五分之三是客户端 | Mini-SWE-Agent's ~100-line loop reports results in the same range as OpenHands' event-sourced engine; roughly three-fifths of OpenCode's non-test source is clients |
| 2 | 厂商原生优化的门槛是谁承担逐 provider 的条件代码成本 | Provider-native optimisations hinge on who pays the per-provider conditional-code cost | Hermes 手写五套传输与 29 个 provider 档案;Pi 把每轮缓存未命中的美元浪费当一等指标;OpenCode 同时发出六家缓存方言 | Hermes hand-rolls five transports and 29 provider profiles; Pi treats per-turn cache-miss dollar waste as a first-class metric; OpenCode emits six cache dialects at once |
| 3 | 提示修辞在经验收敛处收敛,随后随信任校准变薄;十一系统都没有政策级拒绝语言 | Prompt rhetoric converges where experience converges and thins as trust calibrates; no policy-level refusal language anywhere | 「不要镀金」在六个独立 scaffold 里措辞近乎同构;4 月的禁止自动提交规则到 7 月被反转或删除 | Anti-gold-plating language appears near-isomorphically in six independently developed scaffolds; the April no-commit rule was reversed or dropped by July |
| 4 | 文件编辑策略是代码修改准确率的最重要决定因素之一,且按模型多态 | File-editing strategy is a top determinant of accuracy, and is model-aware | Aider 的 prompt-class 工厂;OpenCode 按模型换工具集;Mistral Vibe 一个季度内从模糊匹配转向精确匹配 | Aider's prompt-class factory; OpenCode swapping toolsets per model; Mistral Vibe moving from fuzzy to exact matching within one quarter |
| 5 | 持久记忆取代上下文压缩成为前沿;四类记忆写入路径并存 | Persistent memory replaced compaction as the frontier; four write-path governance models coexist | 压缩已收敛(7/11 阈值触发 LLM 摘要);路径分为 agent 自主维护、人工把关、模型直写但有界、回合前召回 | Compaction has converged (7/11 use threshold-triggered summarisation); paths split into agent-maintained, human-gated, bounded-direct and pre-turn recall |
| 6 | OS 级沙箱昂贵,而且是选择而非规模的必然结果 | OS sandboxing is expensive, and a choice rather than a consequence of scale | 4 月的「规模蕴含沙箱」相关性被打破:Hermes 与 OpenCode 属最大系统却零 OS 级隔离 | April's size-implies-sandbox correlation broke: Hermes and OpenCode are among the largest systems with zero OS-level isolation |
| 7 | coordinator-worker 形态独立涌现;子 agent 协调大多仍在进程内 | Coordinator-worker emerged independently; sub-agent coordination mostly stays in-process | 四个厂商系统全部出现该形态;九分之八的多 agent 系统其子 agent 协调用进程内原语 | All four vendor systems show the shape; eight of nine multi-agent systems coordinate sub-agents with in-process primitives |
| 8 | Skills 超过 MCP 成为采用最广的扩展标准 | Skills overtook MCP as the most-adopted extensibility standard | SKILL.md 9/11 对 MCP 8/11;延迟加载 8/9;出现注册表、信任层级、隔离与 agent 自撰技能 | SKILL.md 9/11 against MCP 8/11; deferred loading 8/9; registries, trust tiers, quarantine and the first agent-authored skills |
| 11 | 协议位置从两角色变三角色:ACP 新增 harness 托管 | Protocol placement went from two roles to three: ACP added harness hosting | ACP 6/11;OpenHands 把 Claude Code、Codex、Gemini CLI 当可互换后端;A2A 仍只有 Gemini CLI | ACP ships in 6/11; OpenHands runs Claude Code, Codex or Gemini CLI as interchangeable backends; A2A remains Gemini CLI's bet alone |
| 12 | 平台化转变已完成:竞争单位从循环移到循环之外的生态面 | The platform turn is complete: the competitive unit moved from the loop to the ecosystem around it | SDK 化发布、框架厂商发 harness、市场与信任层级、跨厂商会话导入、MDM 治理、元 harness 编排十一家厂商 | SDK-shaped releases, framework vendors shipping harnesses, marketplaces and trust tiers, cross-vendor session importers, MDM governance, and a meta-harness orchestrating eleven vendors |
两条经受过复核的缺席
Two absences that survived re-auditing
这两条是全文最可执行的结论,因为它们说的是不要做什么。作者检查了每个依赖清单,并在三种语言里 grep 了十余个通用框架的 import——跨十二棵树复核两次,相隔三个月;语料还经历了一次三倍扩展。结论没有被撼动。
These are the most actionable findings because they are about what not to build. Every dependency manifest was inspected and every tree grepped for imports of a dozen general-purpose frameworks — across twelve trees, twice, three months apart, with the corpus tripling in between. The result held.
作者给出的解释是失败模式的权重:生产 harness 会修改真实代码,因此可调试性与提示透明度压过框架复用带来的便利。他们也把这条结论明确标为「结构上保守」——没有追内部 fork、动态 importlib 与 require 加载的插件、以及转译产物,所以不能说这种用法不存在,只能说没被这轮审计发现。
The explanation given is the weight of the failure mode: production harnesses mutate real code, so debuggability and prompt transparency outweigh the convenience of framework reuse. The authors also label the finding structurally conservative — internal forks, dynamically imported plugins and transpiled distributions were not traced, so the claim is that such use was not found, not that it does not exist.
18 条设计建议与 90 行骨架
18 recommendations and a 90-line scaffold
论文的最后一节是给从业者的清单:每条建议都引用支持它的观察、点名实现它的系统,并指出可能让人做出不同选择的取舍。
The last section is a practitioner's guide: each recommendation cites the observations that support it, names the systems that implement it, and identifies the trade-off that might justify a different choice.
论文还给了一个约 90 行的最小可用 harness,直接实现 18 条中的 10 条、其余兼容。它没有框架依赖、没有 RAG、没有向量库、没有多 agent 编排、也没有沙箱——正好与两条缺席结论一致。作者明确说这是推测(无证明):这个骨架在前沿模型上会接近 Mini-SWE-Agent 的数字,而再往上走主要是模型能力问题,不是 scaffold 问题。
The paper also ships a roughly 90-line minimum-viable harness that realises ten of the eighteen recommendations directly and stays compatible with the rest. It has no framework dependency, no RAG, no vector store, no multi-agent orchestration and no sandbox — exactly in line with the twin absences. The authors are explicit that the follow-on claim is a conjecture, not a proof: that this scaffold would approach Mini-SWE-Agent's numbers on a frontier model, and that moving beyond those numbers is mostly a model-capability question rather than a scaffold question.
90 天纵向与平台化
Ninety days of evolution, and the platform turn
因为 4 月版的八个系统是重新钉版本而非被替换,论文手里有一份受控的纵向样本:同一批 harness 的源码 diff 跨越一个季度。结论是「收敛变成了模仿」——Codex 逐字采用 Claude Code 的 hook 词汇并发布会话与设置导入器;OpenHands 读取 Claude Code 的插件格式;行为政策从提示散文迁移到配置(feature flag 与服务端 A/B 测试)。4 月版的三条观察被新证据实质修订。
Because the eight systems from the April edition were re-pinned rather than replaced, the study holds a controlled longitudinal sample: the same harnesses, source-diffed across one quarter. The result is convergence becoming imitation — Codex adopts Claude Code's hook vocabulary verbatim and ships session and settings importers, OpenHands reads Claude Code's plugin format, and behavioural policy migrates from prompt prose into configuration and server-side A/B tests. Three of the April observations were substantively revised on new evidence.
由此得出的判断是:编码 harness 在 2026 年上半年完成了从工具到平台的转变。证据是具名产物:harness 以可导入 SDK 的形式发布,而框架厂商反过来发布自己的 harness;能力分发有了市场、注册表与信任层级;厂商为彼此的磁盘状态写导入器;出现了 MDM 级别的治理层;agent 本身以 OpenAI 兼容端点暴露;一个元 harness 在一个 API 后面编排十一家厂商的 harness,重新实现其中昂贵的部分、套利专有的部分。作者因此把竞争单位从 agent 循环重新定位到循环之外的生态面。
The resulting judgement is that the coding harness completed its turn from tool to platform in the first half of 2026. The evidence is a list of named artifacts: harnesses ship as importable SDKs while framework vendors ship harnesses; capability distribution gained marketplaces, registries and trust tiers; vendors write importers for each other's on-disk state; MDM-grade governance appeared; the agent itself became addressable behind an OpenAI-compatible endpoint; and one meta-harness orchestrates eleven vendor harnesses behind a single API, re-implementing the expensive parts and arbitraging the proprietary ones. The authors therefore relocate the competitive unit from the agent loop to the ecosystem around it.
批判性评估
Critical assessment
这篇论文的价值在于把「harness 到底是什么」从讨论变成了可核对的清单:七子系统给了设计自查的格子,29 个模式给了实现选项的目录,18 条建议给了起步顺序,两条缺席给了两个常被默认接受的做法以强反例。
The value here is turning "what is a harness" from a discussion into a checkable inventory: seven subsystems give the audit cells, 29 patterns give the option catalogue, 18 recommendations give the starting order, and two absences give strong counterexamples to two widely assumed practices.
一句话:循环的复杂度不预测能力,代码量主要花在循环之外;生产编码 harness 不用通用框架,也不用向量检索查代码——按这三条去检查自己的 harness,比按榜单排名去抄更有效。 In one line: loop sophistication does not predict capability, the mass sits outside the loop, and production coding harnesses use neither a general-purpose framework nor vector retrieval over code — auditing your own harness against those three beats copying a leaderboard.