Most agent-optimisation methods treat the evaluator as a constant: its scores update the agent, while the evaluator itself never moves. DUET points out that this assumption decays as the agent improves — a weak agent fails visibly, a strong one fails inside plausible-looking artifacts. So it makes the grader an optimisation target too, and lets the two evolve in turns.
阅读结论Verdict三条纪律可直接搬 / 成本账没算清Three reusable rules, one missing cost account
交替更新、用分数变化挑任务、编辑预算当学习率——这三条与 benchmark 解耦;但优化过程的开销从未与基线对齐比较。
Alternating updates, selecting tasks by score change, and treating the edit budget as a learning rate are all benchmark-agnostic — but the optimisation cost is never compared against the baselines.
96.5 / 89.4Finance Agent Benchmark 与 τ²-Bench 上参考无关的 solver 分数,均高于使用参考评测器的基线Reference-free solver scores on Finance Agent Benchmark and τ²-Bench, both above baselines that use the benchmark reference evaluator
67.9 → 75.0grader 成对一致率的提升,但四个 benchmark 里只有两个有实质变化Grader pairwise agreement, up from 67.9 — but only two of the four benchmarks moved materially
95.4 / 70.1每轮同时更新两个 agent 的结果,明显差于交替更新(96.5 / 75.0)Updating both agents every round scores clearly worse than alternating (96.5 / 75.0)
论文Paper
DUET: Co-Evolving Solver and Grader Agents
Fengyu Gao(University of Virginia,实习期间完成)、Sourav Pal、Austin Z. Henley、Arjun Radhakrishna、Gustavo Soares(Microsoft)。
Fengyu Gao (University of Virginia, work done during an internship at Microsoft), Sourav Pal, Austin Z. Henley, Arjun Radhakrishna and Gustavo Soares (Microsoft).
Methods such as GEPA and SkillOpt share one loop: run the agent, evaluate the output, update the agent components using the scores or critiques. The paper calls the task-performing side the solver and the evaluating side the grader, and notes that the loop's effectiveness depends directly on the quality of the grader's feedback.
The gap is not that the grader is inaccurate; it is that the grader is treated as static. That assumption breaks as the solver improves: a weak solver fails obviously, while a strong one produces results that look reasonable and only reveal their errors on closer inspection. The original grader misses those new failure modes and finer quality differences, so its feedback carries less information. This bites hardest on open-ended tasks, where several execution paths can be valid and quality depends on the full process and artifacts rather than a single reference answer.
So two questions couple: how should the agent evolve, and how should we evaluate an agent that keeps evolving? The answer here is to evolve both — evaluation stops being a fixed feedback source and becomes a first-class optimisation target.
The formal setup is deliberately narrow: only the solver's component θ and the grader's component φ are optimised (a system prompt, a skill library, or both), while the underlying models, tools and environments stay frozen. Training uses no reference answers and no ground-truth labels; a separate labelled validation set is used only for checkpoint selection and early stopping.
图 1 · 每轮只改一边。奇数轮更新 grader,偶数轮更新 solver,未更新的一侧冻结——这样后面表现的变化才能归因到某一次具体更新。依据原文 Figure 1 与 Algorithm 1 重绘。Figure 1 - One side changes per round. The grader is updated on odd rounds, the solver on even rounds, and the other stays frozen so later changes can be attributed to a specific update. Redrawn from Figure 1 and Algorithm 1.
设计要素
Design element
怎么做
How it works
为什么
Why
自适应选任务
Adaptive task selection
用同一任务相邻两次评测的分数差 r 等于 s(k) 减 s(k−1) 的绝对值衡量训练价值,在滑动窗口内取均值,再叠加 SW-UCB 的探索项挑批;热身阶段先用轮询保证每个任务跑过两次。
Training value is measured by the score change between two consecutive runs of the same task — the absolute difference of s(k) and s(k−1) — averaged over a sliding window and combined with a SW-UCB exploration term; a round-robin warm-up runs every task twice first.
A score that still moves means the task is still exposing new solver behaviour or grader weaknesses; a stable score means it is saturated or currently too hard.
交替更新
Alternating updates
奇数轮只改 grader,偶数轮只改 solver。
Only the grader changes on odd rounds, only the solver on even rounds.
同时改两边会同时改变求解行为和评测策略,无法判断是谁带来的变化。
Changing both at once alters both the solving behaviour and the evaluation policy, so the cause of any change becomes unattributable.
An edit budget acts as a textual learning rate, capping discrete changes per revision and decaying on a cosine schedule (4 down to 2). The update agent is asked to make the smallest evidence-supported revision and to avoid task-specific rules.
早期允许大改以探索新机制,后期只允许可归因的细修,避免过拟合到单个任务。
Broad changes early allow exploration, while late rounds only permit attributable refinements, limiting overfitting to individual tasks.
The update module is itself a tool-using agent: it reads the round evidence and the full history, is asked to verify the grader's feedback against the execution artifacts first, and only then proposes a revision. The paper also has it search the web for relevant current practices — the least reproducible element in the whole pipeline, revisited below.
The four benchmarks span spreadsheet manipulation, economically valuable tasks, multi-turn customer service and retrieval-augmented QA. Their metrics mean different things, so absolute values are not comparable across them.
Benchmark
规模与划分
Size and split
任务类型
Task type
指标
Metric
SpreadsheetBench Verified
400 题(275 单元格级、125 工作表级);60 训练 / 20 验证 / 320 测试
400 tasks (275 cell-level, 125 sheet-level); 60 train / 20 val / 320 test
按自然语言指令修改真实 Excel 工作簿
Editing real Excel workbooks from natural-language instructions
任务准确率(所有被评单元格一致)
Task accuracy (all evaluated cells match)
GDPval
参考交付物含 Excel 的全部 60 题
All 60 tasks whose reference deliverable includes an Excel workbook
经济价值任务
Economically valuable tasks
与人类金标准交付物的成对胜率
Pairwise win rate against the human gold deliverable
τ²-Bench
150 题(航空、零售、电信)
150 tasks (airline, retail, telecom)
有状态环境中的多轮客服,模拟用户与领域工具
Multi-turn customer service in a stateful environment with a simulated user
成功率
Success rate
Finance Agent Benchmark
126 题
126 tasks
需要检索与引用的长文本金融问答
Long-form financial QA requiring retrieval and citations
Both the solver and the grader default to Claude Opus 4.8 across 3 random seeds, with the update module running on GitHub Copilot CLI. Hyperparameters are chosen heuristically per benchmark: batch size 8 / 4 / 8 / 8, SW-UCB window 20 / 8 / 10 / 10, maximum rounds 50 / 40 / 30 / 30, exploration coefficient 0.01 and stopping patience 10 rounds. Everything runs on a single Azure VM (8 physical cores, 64 GB RAM, Windows 11, with a local Excel install for the spreadsheet work); one finance run takes roughly 15 hours.
Grader performance is measured by pairwise agreement: the grader scores each output independently and earns 1 for preferring the reference-preferred output, 0.5 for a tie and 0 for an inversion. The evaluation sets are built accordingly — for Finance and GDPval, multiple candidates are generated per task and ranked by the benchmark reference evaluator, keeping only reliable preferences and further selecting cases where surface cues favour the worse output; for τ² and Spreadsheet, whose native evaluators are binary, one output per task is generated by a different model so each pair has exactly one correct answer.
主结果:solver 在四个 benchmark 上都最强
Results: the strongest solver on all four benchmarks
The starting point is an empty-prompt solver and grader — DUET begins with an evaluator that has no domain knowledge at all.
图 2 · 参考无关设置下的 solver 分数(Claude Opus 4.8,3 seeds 均值)。四个 benchmark 的指标各不相同,只能各自与自己的起点比较。数据来源:原文 Sec.5.1, Table 1。Figure 2 - Solver scores with a reference-free grader (Claude Opus 4.8, 3-seed mean). The four benchmarks use different metrics, so each row should be read against its own starting point rather than across. Source: Sec.5.1, Table 1.
One row is easy to misread. Reference-free DUET beats the baselines that use the benchmark reference evaluator on Finance (96.5 against 95.7 and 95.9) and τ²-Bench (89.4 against 83.9 and 82.0) — the point the authors emphasise. But it sits below its own oracle variant on GDPval (52.2 against 54.9) and ties it on SpreadsheetBench. The accurate reading is that reference-free co-evolution is no worse than baselines given a reference evaluator, not that reference-free is always better.
Two further properties are worth noting. Extending the optimisable target from the system prompt to prompt plus skill library lifts τ²-Bench to 91.9 (from 89.4 with prompt only). And on Claude Sonnet 4.6 the Finance score rises from 89.6 to 91.3, so the direction survives a model swap.
The authors offer an explanation, and it holds up: the native evaluators for τ²-Bench and SpreadsheetBench return only binary success or failure, so the only pairs available are one-correct-one-wrong — a difference the initial grader already recognises. The subtler comparisons that would reveal better judgement never appear in those evaluation sets. The converse also holds: the claim that co-evolution materially improves the evaluator is currently verified on two benchmarks, not four.
All three ablations run on Finance Agent Benchmark. They point the same way, and toward the same conclusion: most of the movement is on the grader side (67.9 up to 75.0), while the solver differences stay small.
图 4 · 三组消融。最反直觉的一行是「固定 solver」:只优化 grader 几乎没带来提升(67.9 到 68.8),说明评测者的进步需要持续进化的求解者提供新挑战。数据来源:原文 Sec.5.3, Figure 3(a)(b)(c),数值按原图读取。Figure 4 - The three ablations. The most counter-intuitive row is the fixed solver: optimising only the grader barely moves it (67.9 to 68.8), implying that evaluator progress needs an evolving solver to keep supplying new challenges. Source: Sec.5.3, Figure 3(a)(b)(c), values read off the figures.
One easily missed detail: in the first ablation, the sequential scheme — optimise the grader first, freeze it, then optimise the solver — already captures most of the solver gain (95.7 against 96.5 for alternation). What alternation adds shows up mainly on the grader side (68.8 to 75.0), a difference the text does not unpack.
批判性评估
Critical assessment
强证据Strong evidence四个跨域 benchmark、三个随机种子、报告均值与标准差,并且同时报告 solver 与 grader 两侧;三组消融方向一致、误差棒可见、都在同一 benchmark 上做;起点与基线对齐(空提示),另外做了 oracle 评测器对照与「固定参考评测器只优化 solver」的对照,把信息条件与方法差异分开;专设对抗性初始化实验,并如实报告对抗起点下 grader 几乎没提升(68.0 到 68.8);连 checkpoint 选择方式也做了消融,而不是只展示最优配置。Four cross-domain benchmarks, three seeds, means with standard deviations, and both solver and grader reported. The three ablations agree in direction, show error bars and run on the same benchmark. The starting point matches the baselines (empty prompt), and separate controls for the oracle evaluator and for fixing the reference grader while optimising only the solver disentangle the information condition from the method. There is a dedicated adversarial-initialisation experiment that honestly reports a near-flat grader result (68.0 to 68.8), and even checkpoint selection gets an ablation instead of only the best configuration being shown.
中等或弱证据Weaker evidence优化过程成本从未与基线对齐比较——成本表只统计最终 agent 的推理开销,而 DUET 有两个优化目标、还要调用更新 agent;grader 的实质提升只在四个 benchmark 中的两个上成立;三组消融只在一个 benchmark 上做;grader 评测集的偏好由 benchmark 参考评测器确定,grader 又被要求去对齐这些偏好,构成一定程度的目标循环,结论应读成「更接近参考评测器」而不是「更接近人类判断」;更新模块会检索网页,既没有做必要性消融,也没有讨论不可复现与测试集泄漏;提示更新的示例是作者挑选的,没有量化任务专属规则的比例;正文写自适应选任务的 grader 一致率为 62.7%,而 Figure 3(b) 标注 62.6%。The optimisation cost is never compared against the baselines: the cost table counts only the final agent's inference spend, while DUET carries two optimisation targets and calls an update agent. The grader gains materialise on two of four benchmarks. All three ablations run on a single benchmark. Grader evaluation preferences are set by the benchmark reference evaluator while the grader is being optimised to match them, a degree of objective circularity — read the result as closer to the reference evaluator, not closer to human judgement. The update module searches the web, with no necessity ablation and no discussion of reproducibility or test-set leakage. The prompt-update examples are author-selected, with no quantification of task-specific rules. And the text reports 62.7% for round-robin grader agreement where Figure 3(b) shows 62.6%.
结论是否超出证据Does the claim exceed the evidence「四个 benchmark 上一致超过固定 grader 的基线」完全成立。「超过使用参考评测器的基线」需要限定到具体 benchmark:成立的是与基线的比较,且在 Finance 与 τ²-Bench 上;GDPval 上它低于自己的 oracle 变体。「可以降低对精心设计初始 grader 的依赖」证据偏中等——空提示起点确实能超过基线,但论文没有对比「精心设计的初始 grader 加 DUET」,而它自己的知情初始化实验里终点反而更低(95.0 对空提示的 96.5)。「不做验证集选择与早停也有相当表现」口径偏宽:solver 相当,但 grader 一致率下降了 4.5 分。"Consistently beats fixed-grader baselines on all four benchmarks" holds. "Exceeds baselines that use the reference evaluator" needs a benchmark qualifier: it holds against the baselines on Finance and τ²-Bench, and on GDPval it sits below its own oracle variant. "Reduces the need for a carefully designed initial grader" is moderate at best — the empty-prompt start does beat the baselines, but the paper never compares against a well-designed initial grader plus DUET, and in its own informed-initialisation experiment the endpoint is lower (95.0 against 96.5). "Comparable performance without validation-based selection" overstates things: the solver is comparable, but grader agreement drops 4.5 points.
This paper lines up with two of my local projects in two different ways: one is about optimising acceptance criteria, the other about giving grading quality its own metric.
本地 Agent Harness:让验收闸门也成为被优化对象
Local Agent Harness: make the acceptance gate an optimisation target
The harness currently writes acceptance criteria into validators and attaches them to task declarations, but those criteria are human-maintained and static. DUET's move is to let the criteria evolve with the thing they judge, under three rules: change only one side per round, pick the tasks still worth re-running by how much their score moved between consecutive runs, and cap the size of each revision. The most directly transferable finding is the counter-intuitive one from the ablation: with the judged object frozen, optimising the judge does almost nothing.
点睛:判分质量需要自己的指标与更新机制
Eyedot: grading quality needs its own metric and update loop
The project's current evaluation covers grading outcomes — per-point accuracy, calibration, needs-review flags — but not whether the grader itself can tell two candidate answers apart. DUET's pairwise agreement (1 for the preferred output, 0.5 for a tie, 0 for an inversion) and its hard-pair construction (generate several candidates, rank them with a reference standard, then keep the pairs where surface cues favour the worse one) transfer directly. And the constraint that the update module makes the smallest evidence-supported revision addresses exactly where rubric iteration goes wrong: ever finer, ever more aligned to the training items.
Both projects share one conclusion: acceptance criteria are not a constant written once, but an object that needs its own evaluation and its own update loop — changing one side at a time, or you will not know what deserves the credit.
DUET's contribution is not another prompt-optimisation trick. It is taking an engineering intuition and making it executable: when the thing being optimised keeps changing, the standard that scores it goes stale.
一句话:把评测者放进优化循环,用两条纪律守住它——每轮只改一边以便归因,只做证据支持的最小修订以便不跑偏;但要自己算一遍优化成本,因为论文没算。In one line: put the evaluator inside the optimisation loop and hold it with two rules — change one side per round so effects stay attributable, and make only the smallest evidence-supported revision so it does not drift; then do the cost arithmetic yourself, because the paper does not.