Paper Reading / Agent Optimisation

DUET 论文精读:把 grader 从固定反馈源变成一等优化目标,用交替更新、自适应选任务和编辑预算三条纪律,在四个 agent benchmark 上同时提升执行与评测。

DUET:评测者一起进化

被优化的东西在变,评价它的标准就会过时。

主流 agent 优化方法都把评测者当成常量:用它的分数去改 agent,自己却一动不动。DUET 指出这个假设会随 agent 变强而失效——弱 agent 的失败很显眼,强 agent 的错藏在看起来合理的产物里。于是它把评测者本身也变成优化目标,让两者轮流进化。

Most agent-optimisation methods treat the evaluator as a constant: its scores update the agent, while the evaluator itself never moves. DUET points out that this assumption decays as the agent improves — a weak agent fails visibly, a strong one fails inside plausible-looking artifacts. So it makes the grader an optimisation target too, and lets the two evolve in turns.

阅读结论 Verdict 三条纪律可直接搬 / 成本账没算清 Three reusable rules, one missing cost account 交替更新、用分数变化挑任务、编辑预算当学习率——这三条与 benchmark 解耦;但优化过程的开销从未与基线对齐比较。 Alternating updates, selecting tasks by score change, and treating the edit budget as a learning rate are all benchmark-agnostic — but the optimisation cost is never compared against the baselines.
96.5 / 89.4 Finance Agent Benchmark 与 τ²-Bench 上参考无关的 solver 分数,均高于使用参考评测器的基线 Reference-free solver scores on Finance Agent Benchmark and τ²-Bench, both above baselines that use the benchmark reference evaluator
67.9 → 75.0 grader 成对一致率的提升,但四个 benchmark 里只有两个有实质变化 Grader pairwise agreement, up from 67.9 — but only two of the four benchmarks moved materially
95.4 / 70.1 每轮同时更新两个 agent 的结果,明显差于交替更新(96.5 / 75.0) Updating both agents every round scores clearly worse than alternating (96.5 / 75.0)
论文 Paper

DUET: Co-Evolving Solver and Grader Agents

Fengyu Gao(University of Virginia,实习期间完成)、Sourav Pal、Austin Z. Henley、Arjun Radhakrishna、Gustavo Soares(Microsoft)。

Fengyu Gao (University of Virginia, work done during an internship at Microsoft), Sourav Pal, Austin Z. Henley, Arjun Radhakrishna and Gustavo Soares (Microsoft).

来源 Source

arXiv:2610.04087v1

2026-10-02 提交。本页按官方 HTML 版精读,Figure 3 的三张消融图为原图读数。

Submitted 2026-10-02. This note reads the official HTML version; the three ablation charts in Figure 3 were read off the original figures.

官方资源 Artifacts

论文未给出代码或数据

No code or data link provided

论文未提供代码仓库或项目页;默认模型为 Claude Opus 4.8,更新模块基于 GitHub Copilot CLI。

The paper provides no repository or project page. The default model is Claude Opus 4.8 and the update module runs on GitHub Copilot CLI.

问题:评测者被当成常量

Problem: the grader treated as a constant

GEPA、SkillOpt 这类方法共用同一套流程:执行 agent、评测输出、用分数或评语更新 agent 组件。论文把执行任务的叫 solver,把评测它的叫 grader,并指出整个循环的效果直接取决于 grader 反馈的质量。

Methods such as GEPA and SkillOpt share one loop: run the agent, evaluate the output, update the agent components using the scores or critiques. The paper calls the task-performing side the solver and the evaluating side the grader, and notes that the loop's effectiveness depends directly on the quality of the grader's feedback.

缺口不是「grader 不够准」,而是grader 被当成静止的。这个假设在 solver 变强之后就不成立了:弱 solver 的失败显而易见,强 solver 会产出看起来合理、需要细看才发现错误的结果。原有 grader 捕捉不到这些新失败模式与更细的质量差异,反馈的信息量随之下降。开放式任务尤其严重——多条执行路径都可能有效,质量取决于完整的执行过程与产物,而不是单一参考答案。

The gap is not that the grader is inaccurate; it is that the grader is treated as static. That assumption breaks as the solver improves: a weak solver fails obviously, while a strong one produces results that look reasonable and only reveal their errors on closer inspection. The original grader misses those new failure modes and finer quality differences, so its feedback carries less information. This bites hardest on open-ended tasks, where several execution paths can be valid and quality depends on the full process and artifacts rather than a single reference answer.

于是两个问题互相耦合:agent 该怎么进化,以及在 agent 不断进化时该怎么评测它。论文的答案是让两者一起进化——评测从固定的反馈源,变成一等优化目标。

So two questions couple: how should the agent evolve, and how should we evaluate an agent that keeps evolving? The answer here is to evolve both — evaluation stops being a fixed feedback source and becomes a first-class optimisation target.

方法:三步闭环与三条纪律

Method: a three-step loop and three rules

形式化很克制:只优化 solver 的组件 θ 与 grader 的组件 φ(系统提示、技能库,或两者),底层模型、工具与执行环境全部冻结。训练不使用参考答案或真值标签;单独的标注验证集只用于 checkpoint 选择与早停。

The formal setup is deliberately narrow: only the solver's component θ and the grader's component φ are optimised (a system prompt, a skill library, or both), while the underlying models, tools and environments stay frozen. Training uses no reference answers and no ground-truth labels; a separate labelled validation set is used only for checkpoint selection and early stopping.

DUET 的三步闭环:自适应选任务、solver 执行与 grader 评测、交替与有界更新 The DUET loop in three steps: adaptive task selection, solver execution and grader evaluation, alternating bounded updates
图 1 · 每轮只改一边。奇数轮更新 grader,偶数轮更新 solver,未更新的一侧冻结——这样后面表现的变化才能归因到某一次具体更新。依据原文 Figure 1 与 Algorithm 1 重绘。
Figure 1 - One side changes per round. The grader is updated on odd rounds, the solver on even rounds, and the other stays frozen so later changes can be attributed to a specific update. Redrawn from Figure 1 and Algorithm 1.
设计要素 Design element 怎么做 How it works 为什么 Why
自适应选任务 Adaptive task selection 用同一任务相邻两次评测的分数差 r 等于 s(k) 减 s(k−1) 的绝对值衡量训练价值,在滑动窗口内取均值,再叠加 SW-UCB 的探索项挑批;热身阶段先用轮询保证每个任务跑过两次。 Training value is measured by the score change between two consecutive runs of the same task — the absolute difference of s(k) and s(k−1) — averaged over a sliding window and combined with a SW-UCB exploration term; a round-robin warm-up runs every task twice first. 分数还在变意味着任务还在暴露新的 solver 行为或 grader 弱点;分数稳定说明它已饱和或当前太难。 A score that still moves means the task is still exposing new solver behaviour or grader weaknesses; a stable score means it is saturated or currently too hard.
交替更新 Alternating updates 奇数轮只改 grader,偶数轮只改 solver。 Only the grader changes on odd rounds, only the solver on even rounds. 同时改两边会同时改变求解行为和评测策略,无法判断是谁带来的变化。 Changing both at once alters both the solving behaviour and the evaluation policy, so the cause of any change becomes unattributable.
有界更新 Bounded update 把编辑预算当文本学习率,限制单次修订的离散改动条数,并随轮次余弦递减(4 降到 2)。更新 agent 被要求只做证据支持的最小修订,并避免任务专属规则。 An edit budget acts as a textual learning rate, capping discrete changes per revision and decaying on a cosine schedule (4 down to 2). The update agent is asked to make the smallest evidence-supported revision and to avoid task-specific rules. 早期允许大改以探索新机制,后期只允许可归因的细修,避免过拟合到单个任务。 Broad changes early allow exploration, while late rounds only permit attributable refinements, limiting overfitting to individual tasks.

更新模块本身也是一个会用工具的 agent:它读本轮证据与全部历史,被要求先核对 grader 的反馈与执行产物是否一致,然后才提出修订。论文还让它检索网页上的相关实践——这是全文最不可复现的一环,后文会展开。

The update module is itself a tool-using agent: it reads the round evidence and the full history, is asked to verify the grader's feedback against the execution artifacts first, and only then proposes a revision. The paper also has it search the web for relevant current practices — the least reproducible element in the whole pipeline, revisited below.

实验设置:四个 benchmark

Setup: four benchmarks

四个 benchmark 横跨表格操作、经济价值任务、多轮客服与检索增强问答。它们的指标语义各不相同,因此不能横向比较绝对值。

The four benchmarks span spreadsheet manipulation, economically valuable tasks, multi-turn customer service and retrieval-augmented QA. Their metrics mean different things, so absolute values are not comparable across them.

Benchmark 规模与划分 Size and split 任务类型 Task type 指标 Metric
SpreadsheetBench Verified 400 题(275 单元格级、125 工作表级);60 训练 / 20 验证 / 320 测试 400 tasks (275 cell-level, 125 sheet-level); 60 train / 20 val / 320 test 按自然语言指令修改真实 Excel 工作簿 Editing real Excel workbooks from natural-language instructions 任务准确率(所有被评单元格一致) Task accuracy (all evaluated cells match)
GDPval 参考交付物含 Excel 的全部 60 题 All 60 tasks whose reference deliverable includes an Excel workbook 经济价值任务 Economically valuable tasks 与人类金标准交付物的成对胜率 Pairwise win rate against the human gold deliverable
τ²-Bench 150 题(航空、零售、电信) 150 tasks (airline, retail, telecom) 有状态环境中的多轮客服,模拟用户与领域工具 Multi-turn customer service in a stateful environment with a simulated user 成功率 Success rate
Finance Agent Benchmark 126 题 126 tasks 需要检索与引用的长文本金融问答 Long-form financial QA requiring retrieval and citations 七个 rubric 维度的评委均分 Mean judge score over seven rubric dimensions

默认 solver 与 grader 都用 Claude Opus 4.8,跑 3 个随机种子;更新模块基于 GitHub Copilot CLI。超参按 benchmark 规模启发式选取:批大小 8 / 4 / 8 / 8,SW-UCB 窗口 20 / 8 / 10 / 10,最大轮数 50 / 40 / 30 / 30,探索系数 0.01,停等耐心 10 轮。全部实验在一台 Azure VM 上完成(8 物理核、64 GB 内存、Windows 11,表格实验依赖本地 Excel),金融 benchmark 跑一次约 15 小时。

Both the solver and the grader default to Claude Opus 4.8 across 3 random seeds, with the update module running on GitHub Copilot CLI. Hyperparameters are chosen heuristically per benchmark: batch size 8 / 4 / 8 / 8, SW-UCB window 20 / 8 / 10 / 10, maximum rounds 50 / 40 / 30 / 30, exploration coefficient 0.01 and stopping patience 10 rounds. Everything runs on a single Azure VM (8 physical cores, 64 GB RAM, Windows 11, with a local Excel install for the spreadsheet work); one finance run takes roughly 15 hours.

grader 的评测方式是成对一致率:grader 独立给两个输出打分,偏向参考更优的输出得 1 分,平局 0.5 分,反向 0 分。评测集这样构造——Finance 与 GDPval 先用 Claude Opus 4.8 为每题生成多个候选,再用 benchmark 参考评测器确定相对质量,保留偏好可靠的对并进一步挑出「表面线索偏向较差输出」的难例;τ² 与 Spreadsheet 的原生评测器只给二值结果,就用不同模型各生成一个输出,保留一对一错的对。

Grader performance is measured by pairwise agreement: the grader scores each output independently and earns 1 for preferring the reference-preferred output, 0.5 for a tie and 0 for an inversion. The evaluation sets are built accordingly — for Finance and GDPval, multiple candidates are generated per task and ranked by the benchmark reference evaluator, keeping only reliable preferences and further selecting cases where surface cues favour the worse output; for τ² and Spreadsheet, whose native evaluators are binary, one output per task is generated by a different model so each pair has exactly one correct answer.

主结果:solver 在四个 benchmark 上都最强

Results: the strongest solver on all four benchmarks

起点是空提示的 solver 与 grader——也就是说,DUET 从一个没有任何领域知识的评测者开始优化。

The starting point is an empty-prompt solver and grader — DUET begins with an evaluator that has no domain knowledge at all.

四个 benchmark 上未优化起点、SkillOpt、GEPA 与 DUET 的 solver 分数,以及相对起点的提升 Solver scores for the unoptimised start, SkillOpt, GEPA and DUET across four benchmarks, plus the gain over the start
图 2 · 参考无关设置下的 solver 分数(Claude Opus 4.8,3 seeds 均值)。四个 benchmark 的指标各不相同,只能各自与自己的起点比较。数据来源:原文 Sec.5.1, Table 1。
Figure 2 - Solver scores with a reference-free grader (Claude Opus 4.8, 3-seed mean). The four benchmarks use different metrics, so each row should be read against its own starting point rather than across. Source: Sec.5.1, Table 1.
方法 Method Finance τ²-Bench GDPval SpreadsheetBench
未优化起点Unoptimised A092.880.347.672.6
SkillOpt94.084.748.881.6
GEPA94.483.350.782.8
DUET96.589.452.284.7
SkillOpt(参考评测器)95.783.947.183.4
GEPA(参考评测器)95.982.050.783.0
DUET(参考评测器)96.086.154.984.7

表里有个容易被读错的地方。参考无关的 DUET 在 Finance(96.5)与 τ²-Bench(89.4)上超过了使用 benchmark 参考评测器的基线(95.7 与 95.9、83.9 与 82.0),这正是作者强调的点。但它在 GDPval 上低于自己的 oracle 变体(52.2 对 54.9),在 SpreadsheetBench 上与 oracle 变体打平。所以准确的表述是「参考无关的共同进化不逊于使用参考评测器的基线」,而不是「参考无关总是更好」。

One row is easy to misread. Reference-free DUET beats the baselines that use the benchmark reference evaluator on Finance (96.5 against 95.7 and 95.9) and τ²-Bench (89.4 against 83.9 and 82.0) — the point the authors emphasise. But it sits below its own oracle variant on GDPval (52.2 against 54.9) and ties it on SpreadsheetBench. The accurate reading is that reference-free co-evolution is no worse than baselines given a reference evaluator, not that reference-free is always better.

另外两项属性值得记下:把优化目标从系统提示扩展到提示与技能库之后,τ²-Bench 上进一步提升到 91.9(仅提示时为 89.4);换成 Claude Sonnet 4.6 后 Finance 上从 89.6 升到 91.3,结论方向不变。

Two further properties are worth noting. Extending the optimisable target from the system prompt to prompt plus skill library lifts τ²-Bench to 91.9 (from 89.4 with prompt only). And on Claude Sonnet 4.6 the Finance score rises from 89.6 to 91.3, so the direction survives a model swap.

评测者真的变好了吗

Did the grader actually improve

这是全文最需要仔细读的一节。grader 的成对一致率在四个 benchmark 上的变化是这样的:

This is the section that most rewards careful reading. Here is how grader pairwise agreement moves across the four benchmarks:

四个 benchmark 上 grader 初始与训练后的成对一致率对比 Initial versus post-training grader pairwise agreement across four benchmarks
图 3 · 提升集中在两个 benchmark:Finance 增加 7.1 分、GDPval 增加 10.1 分;τ²-Bench 只有 0.6 分、SpreadsheetBench 只有 0.9 分。数据来源:原文 Sec.5.1, Table 1。
Figure 3 - The gains concentrate in two benchmarks: Finance +7.1 and GDPval +10.1, against only +0.6 on τ²-Bench and +0.9 on SpreadsheetBench. Source: Sec.5.1, Table 1.

作者自己给出了解释,而且解释是合理的:τ²-Bench 与 SpreadsheetBench 的原生评测器只给二值成败,因此只能构造「一对一错」的偏好对,初始 grader 已经能识别这类粗粒度差异,真正能体现判断力提升的细分比较根本没有出现在评测集里。反过来说也成立——「共同进化能显著提升评测者」这一结论,目前只在两个 benchmark 上被验证。

The authors offer an explanation, and it holds up: the native evaluators for τ²-Bench and SpreadsheetBench return only binary success or failure, so the only pairs available are one-correct-one-wrong — a difference the initial grader already recognises. The subtler comparisons that would reveal better judgement never appear in those evaluation sets. The converse also holds: the claim that co-evolution materially improves the evaluator is currently verified on two benchmarks, not four.

消融:三组对照

Ablations: three controls

三组消融都在 Finance Agent Benchmark 上做,方向一致,而且都指向同一个结论:收益主要来自 grader 侧(67.9 升到 75.0),solver 侧的差异反而很小。

All three ablations run on Finance Agent Benchmark. They point the same way, and toward the same conclusion: most of the movement is on the grader side (67.9 up to 75.0), while the solver differences stay small.

Finance Agent Benchmark 上三组消融的 solver 分数与 grader 成对一致率 Solver score and grader pairwise agreement for three ablations on Finance Agent Benchmark
图 4 · 三组消融。最反直觉的一行是「固定 solver」:只优化 grader 几乎没带来提升(67.9 到 68.8),说明评测者的进步需要持续进化的求解者提供新挑战。数据来源:原文 Sec.5.3, Figure 3(a)(b)(c),数值按原图读取。
Figure 4 - The three ablations. The most counter-intuitive row is the fixed solver: optimising only the grader barely moves it (67.9 to 68.8), implying that evaluator progress needs an evolving solver to keep supplying new challenges. Source: Sec.5.3, Figure 3(a)(b)(c), values read off the figures.

还有一个容易被忽略的细节:第一组消融里,「先优化 grader 再冻结、然后优化 solver」这种顺序式做法已经拿到了大部分 solver 收益(95.7,对比交替的 96.5),交替带来的增量主要体现在 grader 侧(68.8 到 75.0)。正文没有对这个差别展开。

One easily missed detail: in the first ablation, the sequential scheme — optimise the grader first, freeze it, then optimise the solver — already captures most of the solver gain (95.7 against 96.5 for alternation). What alternation adds shows up mainly on the grader side (68.8 to 75.0), a difference the text does not unpack.

批判性评估

Critical assessment

强证据 Strong evidence 四个跨域 benchmark、三个随机种子、报告均值与标准差,并且同时报告 solver 与 grader 两侧;三组消融方向一致、误差棒可见、都在同一 benchmark 上做;起点与基线对齐(空提示),另外做了 oracle 评测器对照与「固定参考评测器只优化 solver」的对照,把信息条件与方法差异分开;专设对抗性初始化实验,并如实报告对抗起点下 grader 几乎没提升(68.0 到 68.8);连 checkpoint 选择方式也做了消融,而不是只展示最优配置。 Four cross-domain benchmarks, three seeds, means with standard deviations, and both solver and grader reported. The three ablations agree in direction, show error bars and run on the same benchmark. The starting point matches the baselines (empty prompt), and separate controls for the oracle evaluator and for fixing the reference grader while optimising only the solver disentangle the information condition from the method. There is a dedicated adversarial-initialisation experiment that honestly reports a near-flat grader result (68.0 to 68.8), and even checkpoint selection gets an ablation instead of only the best configuration being shown.
中等或弱证据 Weaker evidence 优化过程成本从未与基线对齐比较——成本表只统计最终 agent 的推理开销,而 DUET 有两个优化目标、还要调用更新 agent;grader 的实质提升只在四个 benchmark 中的两个上成立;三组消融只在一个 benchmark 上做;grader 评测集的偏好由 benchmark 参考评测器确定,grader 又被要求去对齐这些偏好,构成一定程度的目标循环,结论应读成「更接近参考评测器」而不是「更接近人类判断」;更新模块会检索网页,既没有做必要性消融,也没有讨论不可复现与测试集泄漏;提示更新的示例是作者挑选的,没有量化任务专属规则的比例;正文写自适应选任务的 grader 一致率为 62.7%,而 Figure 3(b) 标注 62.6%。 The optimisation cost is never compared against the baselines: the cost table counts only the final agent's inference spend, while DUET carries two optimisation targets and calls an update agent. The grader gains materialise on two of four benchmarks. All three ablations run on a single benchmark. Grader evaluation preferences are set by the benchmark reference evaluator while the grader is being optimised to match them, a degree of objective circularity — read the result as closer to the reference evaluator, not closer to human judgement. The update module searches the web, with no necessity ablation and no discussion of reproducibility or test-set leakage. The prompt-update examples are author-selected, with no quantification of task-specific rules. And the text reports 62.7% for round-robin grader agreement where Figure 3(b) shows 62.6%.
结论是否超出证据 Does the claim exceed the evidence 「四个 benchmark 上一致超过固定 grader 的基线」完全成立。「超过使用参考评测器的基线」需要限定到具体 benchmark:成立的是与基线的比较,且在 Finance 与 τ²-Bench 上;GDPval 上它低于自己的 oracle 变体。「可以降低对精心设计初始 grader 的依赖」证据偏中等——空提示起点确实能超过基线,但论文没有对比「精心设计的初始 grader 加 DUET」,而它自己的知情初始化实验里终点反而更低(95.0 对空提示的 96.5)。「不做验证集选择与早停也有相当表现」口径偏宽:solver 相当,但 grader 一致率下降了 4.5 分。 "Consistently beats fixed-grader baselines on all four benchmarks" holds. "Exceeds baselines that use the reference evaluator" needs a benchmark qualifier: it holds against the baselines on Finance and τ²-Bench, and on GDPval it sits below its own oracle variant. "Reduces the need for a carefully designed initial grader" is moderate at best — the empty-prompt start does beat the baselines, but the paper never compares against a well-designed initial grader plus DUET, and in its own informed-initialisation experiment the endpoint is lower (95.0 against 96.5). "Comparable performance without validation-based selection" overstates things: the solver is comparable, but grader agreement drops 4.5 points.

对本机两个项目的意义

What it means for two local projects

这篇论文与我本机的两个项目各有一处直接对应:一个是「验收标准应该被优化」,一个是「判分质量需要自己的指标」。

This paper lines up with two of my local projects in two different ways: one is about optimising acceptance criteria, the other about giving grading quality its own metric.

本地 Agent Harness:让验收闸门也成为被优化对象

Local Agent Harness: make the acceptance gate an optimisation target

harness 目前把验收标准写进 validators 并挂到任务声明上,但它由人工维护、是静态的。DUET 给的做法是让验收标准随被验收对象一起演化,并用三条纪律控制它:每次只改一边(先改闸门还是先改任务要明确)、用「同一任务相邻两次评测的分数变化」挑出还值得重跑的任务、限制单次修订的规模。最值得直接借用的是消融里那条反直觉结论——固定被评对象时,优化评测者几乎没有用,两者必须一起动。

The harness currently writes acceptance criteria into validators and attaches them to task declarations, but those criteria are human-maintained and static. DUET's move is to let the criteria evolve with the thing they judge, under three rules: change only one side per round, pick the tasks still worth re-running by how much their score moved between consecutive runs, and cap the size of each revision. The most directly transferable finding is the counter-intuitive one from the ablation: with the judged object frozen, optimising the judge does almost nothing.

点睛:判分质量需要自己的指标与更新机制

Eyedot: grading quality needs its own metric and update loop

项目现在的评测只覆盖判分结果(逐点准确率、校准、待复核),没有评测「判分器本身在成对比较上有没有分辨力」。DUET 的成对一致率(胜 1、平 0.5、反向 0)与难例构造方式(多生成几个候选、用参考标准定相对质量、专挑表面线索偏向较差输出的对)可以直接搬过来;而「更新模块只做证据支持的最小修订」这条约束,正好对应评分标准迭代时最容易失控的地方——越改越细、越改越像在背题。

The project's current evaluation covers grading outcomes — per-point accuracy, calibration, needs-review flags — but not whether the grader itself can tell two candidate answers apart. DUET's pairwise agreement (1 for the preferred output, 0.5 for a tie, 0 for an inversion) and its hard-pair construction (generate several candidates, rank them with a reference standard, then keep the pairs where surface cues favour the worse one) transfer directly. And the constraint that the update module makes the smallest evidence-supported revision addresses exactly where rubric iteration goes wrong: ever finer, ever more aligned to the training items.

两个项目共用一条结论:验收标准不是一次写好的常量,而是一个需要被评测、也需要被更新的对象——但每次只改一边,否则改好了也不知道是谁的功劳。

Both projects share one conclusion: acceptance criteria are not a constant written once, but an object that needs its own evaluation and its own update loop — changing one side at a time, or you will not know what deserves the credit.

DUET 的贡献不是又一种提示优化技巧,而是把一句工程直觉落成了可执行流程:当被优化的对象在变,评价它的标准也会过时。

DUET's contribution is not another prompt-optimisation trick. It is taking an engineering intuition and making it executable: when the thing being optimised keeps changing, the standard that scores it goes stale.

一句话:把评测者放进优化循环,用两条纪律守住它——每轮只改一边以便归因,只做证据支持的最小修订以便不跑偏;但要自己算一遍优化成本,因为论文没算。 In one line: put the evaluator inside the optimisation loop and hold it with two rules — change one side per round so effects stay attributable, and make only the smallest evidence-supported revision so it does not drift; then do the cost arithmetic yourself, because the paper does not.