Paper Reading / Self-Evolution

论文精读:自演化搜索 agent 的共谋作弊——出题者与解题者逐渐在同一批错误上达成一致,内部奖励上升而外部正确率不动;CrossFit 用源级交叉拟合反馈把这条回流路径切断。

自演化的共谋作弊

自己出题、自己判卷、再按「判得有多一致」发奖励,系统就会一起错。

自演化 agent 不需要人工出题:proposer 从源文档生成问题与伪标签,solver 作答,再用 solver 在新提案上的表现给 proposer 发奖励。论文指出这条闭环里多了一条回流路径——错误的伪标签训练出会重复该错误的解题者,而这份「一致」又被当成进步回报给出题者。

Self-evolving agents do not need human-written questions: a proposer turns source documents into questions and pseudo-labels, a solver answers them, and the solver's performance on new proposals pays the proposer. The paper shows this loop carries an extra return path — a wrong pseudo-label trains a solver that repeats the error, and that agreement is then rewarded as progress.

阅读结论 Verdict 诊断扎实 / 外部锚点仍是模型 Strong diagnosis, model-based ground truth 机制消融与固定回放做得很干净;但「外部正确性」由 LLM 审计者判定,且主表无方差、训练预算增加七成以上。 The mechanism ablations and fixed-bank replay are clean; but external correctness is judged by an LLM, the main table has no variance, and training cost rises by more than 70%.
0.004 → 0.061 假一致质量随轮次上升(Qwen3.5-4B 的第 1 轮到第 3 轮;9B 从 0.003 升到 0.088) False-agreement mass grows by round for Qwen3.5-4B, from 0.004 to 0.061 (and 0.003 to 0.088 at 9B)
0.7 对 8.8 只做标签验证(MSV)与改反馈来源(CrossFit)带来的下游平均提升,单位是百分点 Downstream average gain in points from verifying labels (MSV) versus changing feedback provenance (CrossFit)
0.710 → 0.679 CrossFit 反而让循环内一致率下降,却把真实标签正确率从 0.747 提到 0.819 CrossFit lowers in-loop agreement while raising true label correctness from 0.747 to 0.819
论文 Paper

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Meijia Chen、Hao Li、Zheng Lu(共同一作)等 15 位作者;Rutgers、UCSD、Michigan、McGill、KFUPM 与独立研究者。

Meijia Chen, Hao Li and Zheng Lu (equal contribution) with 12 co-authors, from Rutgers, UCSD, Michigan, McGill, KFUPM and independent researchers.

来源 Source

arXiv:2609.39102v2

2026-09-30 提交、2026-10-03 修订,21 页。本页按官方 HTML 版精读。

Submitted 2026-09-30, revised 2026-10-03, 21 pages. This note reads the official HTML version.

官方资源 Artifacts

论文未给出代码或数据

No code or data provided

复现依赖 Dr. Zero 的实现与检查点;审计依赖具体商业模型(gpt-6-astra/high)。

Reproduction depends on the Dr. Zero implementation and checkpoints; the audit depends on a specific commercial model (gpt-6-astra/high).

问题:闭环里的「一致陷阱」

Problem: the agreement trap

自演化搜索 agent 的闭环是:proposer 把源文档变成问题与伪标签,准入的样本训练 solver,而 solver 在新提案上的表现决定 proposer 的奖励。循环往复会让提案逐渐逼近 solver 的能力前沿,从而形成无需人工标注的自动课程。

The loop is: a proposer turns source documents into questions and pseudo-labels, admitted samples train the solver, and the solver's performance on new proposals determines the proposer's reward. Repeating this pushes proposals toward the solver's capability frontier, producing an automated curriculum with no human labels.

问题在于,这条闭环让一致率变成了正确率的内生代理:一条错误的伪标签可以训练 solver 在该来源的后续问题上重复同一错误,而这份一致性又被回报给 proposer,强化了下一轮课程里的同一个错误。论文把由此产生的失效模式命名为 co-cheating(共谋作弊):出题者与解题者逐渐在同一批错误上达成一致,内部奖励上升,外部正确率却不涨甚至下降。

The problem is that the loop turns agreement into an endogenous proxy for correctness: a wrong pseudo-label can train the solver to repeat that error on later questions from the same source, and the resulting agreement is paid back to the proposer, reinforcing the same error in the next round's curriculum. The paper names this failure co-cheating: proposer and solver converge on the same errors, so internal reward rises while external correctness does not.

需要说清的是:这里的「共谋」是一个统计现象,作者自己补了一句「不代表有意协调」。引用时宜按「共享错误的自增强」理解,而不是把模型拟人化。

One clarification: co-cheating is a statistical outcome, and the authors themselves note it does not imply intentional coordination. Read it as a self-reinforcing agreement on shared errors rather than anthropomorphising the models.

诊断:把真一致与假一致分开

Diagnosis: separating real from false agreement

论文的做法是加一个完全不参与训练的事后审计:在每一步保存源文档、采纳的伪标签与用于 proposer 奖励的那 5 条 solver 回答,再由一个独立裁判从源文档构造基于证据的参考来判定。审计器不影响准入、模型更新或奖励,因此它测的是训练信号背后的原始样本,而不是循环自己的输出。

The method adds a post-hoc audit that never touches training: at every scheduled step it saves the source document, the adopted pseudo-label and the five solver responses used for proposer reward, and an independent judge builds an evidence-backed reference from the source to score them. Because the auditor affects neither admission, updates nor reward, it measures the samples behind the training signal rather than the loop's own output.

自演化闭环与回流路径,以及循环一致率、假一致质量、应得分数被剥夺三个指标 The self-evolution loop with its return path, plus agreement, false-agreement mass and withheld-credit metrics
图 1 · 闭环里多了一条回流路径。只看循环内一致率 A 会被骗过:一致率上升不等于正确率上升,必须把假一致质量 F 与「应得分数被剥夺」L 一起看。依据原文 Figure 1、Sec.2 与 Table 4 重绘。
Figure 1 - The loop carries an extra return path. Watching in-loop agreement A alone is misleading: rising agreement is not rising correctness, so false-agreement mass F and withheld credit L must be read alongside it. Redrawn from Figure 1, Sec.2 and Table 4.
指标 Metric 含义 Meaning
A 循环观测到的一致率 Agreement as the loop sees it 伪标签与 solver 回答匹配的比例——这正是进入奖励的那个信号 The share of label-response pairs that match — the signal that actually enters the reward
T_P 采纳标签正确率 Adopted-label truth 由审计器对照来源判定标签本身是否成立 Whether the label itself holds up against the source, per the auditor
T_S solver 回答正确率 Solver-response truth 解除标签噪声后的真实解题能力 Real solving ability with label noise removed
F F 假一致质量 False-agreement mass 在同一错误答案上达成一致的比例——本文的核心指标 The share of pairs agreeing on the same incorrect answer — the core metric here
L 应得分数被剥夺 Withheld credit 回答正确却因标签错误而拿不到分数的比例 Correct responses denied credit because the label was wrong
J/E 审计覆盖率 Audit coverage 审计器能给出参考的比例;判不了的样本保持未解决 The share the auditor could resolve; unresolved items stay unresolved

判据很简单:只有当一致率 A 上升、同时 T_P 与 T_S 也上升、且 F 保持低位时,一致率才是可靠的信号。

The reading rule is simple: agreement is only a trustworthy signal when A rises while T_P and T_S also rise and F stays low.

两种干预与各自的作用点

Two interventions at two different points

论文比较了两条路线,它们作用在闭环的不同位置:MSV 管「哪些样本能进训练集」,CrossFit 管「谁来给 proposer 反馈」。

The paper compares two routes that act at different points of the loop: MSV decides which samples enter training, CrossFit decides who supplies the proposer's feedback.

MSV 在准入时做多样本验证,CrossFit 在打分时用另一折训练的辅助 solver 提供反馈 MSV verifies multi-sample consistency at admission, while CrossFit scores with an auxiliary solver trained on the other fold
图 2 · 两种干预。MSV 用同一模型带源采 3 次、不带源采 3 次,两者兼容才准入;CrossFit 把源文档按血缘分成两折,让每个问题都由没见过该来源伪标签的 solver 打分。数据来源:原文 Sec.3.1、Sec.3.2、Figure 2 与 Table 3。
Figure 2 - The two interventions. MSV samples the same model three times with the source and three times without, admitting only compatible majorities; CrossFit splits source documents into two folds so every question is scored by a solver that never saw that source's pseudo-labels. Source: Sec.3.1, Sec.3.2, Figure 2 and Table 3.

CrossFit 的关键设计有三点。第一,划分必须在源文档级别:同一文档派生的问题若分处两折,仍然保留着要消除的复用路径。第二,主 solver 的训练规则完全不变——它依然用全部已准入问题训练,只有塑造 proposer 课程的反馈被替换,因此这是最干净的单变量干预。第三,作者明确它不是把辅助 solver 当成真值判官:两个 solver 仍可能共享预训练错误与重叠证据,它阻断的只是「同源错误被直接回报成奖励」这一条路径。

Three design points matter in CrossFit. First, the split must happen at the source-document level: if questions derived from one document land in both folds, the reuse path survives. Second, the main solver's update rule is untouched — it still trains on every admitted question from both folds, and only the feedback shaping the proposer's curriculum is swapped, which makes this a clean single-variable intervention. Third, the authors state plainly that this does not turn the auxiliary solver into a truth oracle: the two solvers can still share pretraining errors and overlapping evidence; what is blocked is the specific path where a same-source error returns as reward.

主结果:七个搜索 benchmark

Results: seven search benchmarks

评测集固定为 1325 题(NQ、TriviaQA、PopQA、HotpotQA、2WikiMQA、MuSiQue 各 200,Bamboogle 125),每个问题一条贪心搜索轨迹,工具预算与答案抽取一致;所有自演化处理都从同一个公共 checkpoint 出发,且不使用任何人工标注的 QA 训练数据。

The evaluation set is fixed at 1,325 questions (200 each from NQ, TriviaQA, PopQA, HotpotQA, 2WikiMQA and MuSiQue, plus all 125 Bamboogle items), with one greedy search trajectory per question and identical tool budgets and answer extraction. Every self-evolution treatment starts from the same public checkpoint and uses no human-annotated QA training data.

七 benchmark 宏平均 Cover-EM 对比,以及 CrossFit 相对 Dr. Zero 的逐 benchmark 差值 Macro-average Cover-EM across seven benchmarks, plus the per-benchmark difference of CrossFit over Dr. Zero
图 3 · 下游结果。宏平均从 Dr. Zero 的 0.400 / 0.428 提到 CrossFit 的 0.488 / 0.512,即提升 8.8 / 8.4 个百分点;多跳任务的增益(10.0 / 10.9)明显大于单跳(7.3 / 5.2)。数据来源:原文 Sec.4.2, Table 1。
Figure 3 - Downstream results. The macro average rises from 0.400 / 0.428 for Dr. Zero to 0.488 / 0.512 for CrossFit, a gain of 8.8 / 8.4 points; multi-hop tasks gain more (10.0 / 10.9) than single-hop ones (7.3 / 5.2). Source: Sec.4.2, Table 1.
方法 Method Qwen3.5-4B Qwen3.5-9B
起点(不训练)Base0.3840.409
Prompting0.3040.332
R1-Instruct0.3490.379
Search-R10.4010.434
Dr. Zero(耦合闭环)0.4000.428
MSV0.4070.436
CrossFit0.4880.512
MSV+CrossFit0.4910.515

两个对比值得单独强调。第一,只做验证远远不够:MSV 相对 Dr. Zero 只提升 0.7 到 0.8 个百分点,叠加 CrossFit 之后也只比单用 CrossFit 多 0.3 个百分点——改变反馈血缘比改善伪标签质量重要得多。第二,自演化本身并不自动带来收益:耦合闭环(0.400 / 0.428)与外部监督的 Search-R1(0.401 / 0.434)几乎持平,真正的跃升来自修好这条回流路径。

Two comparisons stand out. First, verification alone is far from enough: MSV adds only 0.7 to 0.8 points over Dr. Zero, and combining it with CrossFit adds just 0.3 points beyond CrossFit alone — changing feedback provenance matters much more than improving pseudo-label quality. Second, self-evolution does not pay off by itself: the coupled loop (0.400 / 0.428) sits level with externally supervised Search-R1 (0.401 / 0.434), and the real jump comes from repairing the return path.

机制消融:为什么必须按来源切

Ablations: why the split must follow sources

这是全文方法上最扎实的一节。作者构造了一个固定回放库(3000 条已保存的问题与标签),冻结题库、标签、答案匹配器与评测流程,只改变反馈 solver 的训练来源——这样就能把「机制效果」与「课程变化」分开。

This is the most solid part of the paper. The authors build a fixed replay bank of 3,000 saved questions and labels, then freeze the bank, the labels, the answer matcher and the evaluation procedure while varying only the training provenance of the feedback solver — which separates the mechanism from any change in the curriculum.

对照 Control 目的 Purpose 假一致质量(4B / 9B) False agreement (4B / 9B)
同源辅助 solver Same-source auxiliary 排除「只是换了个模型」 Rule out "just another model" 0.064 / 0.087
全量数据辅助 solver Full-data auxiliary 排除「多训练了一份」 Rule out "one extra training run" 0.058 / 0.069
问题级随机划分 Question-level random split 检验划分粒度 Test the split granularity 0.050 / 0.062
源级划分(主方法) Source-level split (main) 按数据血缘排除 Exclusion by data ancestry 0.004 / 0.001
半预算版本 Half-budget variant 排除「多更新」解释 Rule out "more updates" 0.005 / 0.002

结论很干净:仅仅「另起一个评测模型」不够,按问题随机划分也不够,关键是评测者没见过被评来源派生的伪标签。源级划分还把固定库上的 solver 正确率从 0.687 提到 0.770(4B)、从 0.717 提到 0.868(9B),回放准确率从 88.1% 提到 91.5%、87.0% 提到 91.7%;半预算对照几乎持平,说明增益不是靠更多更新换来的。

The conclusion is clean: adding a second evaluator model is not enough, and neither is a question-level random split; what matters is that the evaluator never saw pseudo-labels derived from the source it judges. The source-level split also raises fixed-bank solver truth from 0.687 to 0.770 (4B) and from 0.717 to 0.868 (9B), with replay accuracy moving from 88.1% to 91.5% and 87.0% to 91.7%; the half-budget control matches this, so the gain does not come from more updates.

还有一组数字值得单独记下:第 3 轮审计里,CrossFit 把假一致质量从 0.061 / 0.088 降到 0.030 / 0.037,把采纳标签正确率从 0.747 / 0.737 提到 0.819 / 0.851。但它同时把循环内一致率从 0.710 / 0.752 降到 0.679 / 0.709,并把「正确回答被误判」的比例抬高。只看内部指标的运维者会认为这个方法在变差——这正是全文最实用的一条工程结论。

One further set of numbers deserves recording: in the round-3 audit CrossFit cuts false agreement from 0.061 / 0.088 to 0.030 / 0.037 and lifts adopted-label truth from 0.747 / 0.737 to 0.819 / 0.851. But it also lowers in-loop agreement from 0.710 / 0.752 to 0.679 / 0.709 and raises the share of correct answers wrongly denied credit. An operator watching only internal metrics would call this a regression — which is the most practical finding in the paper.

成本

Cost

处理 Treatment H200 小时(4B / 9B) 输入 token(百万) Input tokens (M) 裁判请求(千) Judge requests (k) 墙钟(小时) Wall clock (h)
Dr. Zero379 / 47669660.447 / 60
MSV719 / 9031,811296.990 / 113
CrossFit(每折 25 次)CrossFit (25 per fold)650 / 8541,081113.881 / 107
CrossFit(每轮 25 次总计)CrossFit (25 total)515 / 66588987.164 / 83
MSV+CrossFit990 / 1,2812,196350.3124 / 160

CrossFit 增加约 72% 与 79% 的训练预算,MSV 增加约 90%,两者叠加是 2.6 到 2.7 倍;但把辅助预算减半(每轮 25 次总计)可以把增量压到 36% 与 40%,而假一致质量几乎不变(0.005 对 0.004、0.002 对 0.001)。论文没有给出性价比曲线,也没有与「直接用外部监督数据」的成本对比;H200 小时是预留预算(含等待时间),外部审计服务的 GPU 消耗未计入。

CrossFit adds roughly 72% and 79% to the training budget, MSV about 90%, and the combination 2.6 to 2.7 times. Halving the auxiliary budget (25 total updates per round) cuts the overhead to 36% and 40% with almost identical false agreement (0.005 versus 0.004, 0.002 versus 0.001). The paper gives no cost-effectiveness curve and no comparison against simply using externally supervised data; the H200-hours are reserved budgets including waiting time, and the external audit service's GPU usage is not counted.

批判性评估

Critical assessment

强证据 Strong evidence 诊断与训练完全解耦:审计器不参与准入、更新或奖励,测的是训练信号背后的原始样本。机制消融给出三个失败对照(同源、全量、问题级随机划分)与一个成功对照,把「只是多了一个模型」的解释排除干净。固定回放库把反馈来源与课程变化分离,并带 5 seeds 方差,是全文最扎实的一环。七个 benchmark 增益全为正、两种规模方向一致、多跳任务受益更明显。作者还主动写出最危险的反解释(假一致下降可能来自拒绝难任务),并给出配套指标。 The diagnosis is fully decoupled from training: the auditor never touches admission, updates or reward, so it measures the samples behind the training signal. The mechanism ablation supplies three failing controls (same-source, full-data, question-level random split) and one succeeding control, ruling out the explanation that it was merely an extra model. The fixed replay bank separates feedback provenance from curriculum change and carries five-seed variance, the most solid element in the paper. Gains are positive on all seven benchmarks, consistent across two scales, and larger on multi-hop tasks. The authors also volunteer the most dangerous counter-explanation, that lower false agreement may come from rejecting harder tasks, and pair it with the metrics needed to check.
中等或弱证据 Weaker evidence 外部正确性的锚点仍是 LLM:审计由 gpt-6-astra/high 构造参考并判定,全文未报告人工验证样本,而核心指标 F 的可靠性正建立在审计者之上。主表每格是单次运行,没有方差或置信区间(只有固定回放消融给了 5 seeds)。只有 Qwen3.5 一个模型家族的两个规模、三轮演化,无第三方复现。成本没有曲线化,也没有与外部监督路线的总成本对比。审计覆盖率约 85% 到 86%,未解决样本不计入正确率,若其系统性偏向难例则 T_P 会被高估。对「相连来源」(同一网页集群、互相引用的文档)的稳健性未验证,作者自己列为下一步。 External correctness still rests on an LLM: the audit is built and judged by gpt-6-astra/high, no human validation sample is reported, and the reliability of the core metric F depends on that auditor. Each main-table cell is a single run with no variance or interval, and only the replay ablation reports five seeds. The study covers one model family at two scales over three rounds with no third-party replication. Cost is never turned into a curve, nor compared against the total cost of an externally supervised route. Audit coverage is about 85 to 86 percent and unresolved items are excluded from correctness, so a systematic bias toward hard items would inflate T_P. Robustness to connected sources, meaning the same page cluster or mutually citing documents, is untested and named by the authors as future work.
结论是否超出证据 Does the claim exceed the evidence 「共谋作弊存在且随轮次加深」成立;「CrossFit 比只优化伪标签质量更有效」成立(MSV 只提升 0.7 到 0.8,CrossFit 提升 8.8 与 8.4);「增益来自反馈血缘而非课程变化」也成立,固定回放库是最直接的证据——同库同标签,仅换反馈来源就把 F 从 0.058 与 0.073 降到 0.004 与 0.001。作者对「扩展到相连来源」与「更低的端到端成本」明确标为尚未建立,措辞克制。唯一需要限定的是术语:co-cheating 描述的是统计现象,作者自己也补充了「不代表有意协调」。 The existence of co-cheating and its growth by round hold. So does the claim that CrossFit beats merely improving pseudo-label quality, since MSV adds 0.7 to 0.8 points while CrossFit adds 8.8 and 8.4. The attribution claim also holds, with the fixed replay bank as the cleanest evidence: same bank, same labels, only the feedback source changed, and F falls from 0.058 and 0.073 to 0.004 and 0.001. The authors explicitly mark extension to connected sources and lower end-to-end cost as not yet established, which is restrained. The one thing to qualify is the term itself: co-cheating describes a statistical outcome, and the authors note it does not imply intentional coordination.

对本机两个项目的意义

What it means for two local projects

这篇论文与我本机的两个项目都直接相关,而且指向同一件事:当系统既是数据生产者又是评测者时,必须审计评测者自己的血缘。

This paper relates directly to both local projects, and both point at the same thing: when a system both produces the data and judges it, the provenance of the judge itself has to be audited.

本地 Agent Harness:给验收闸门装上血缘字段

Local Agent Harness: give the acceptance gate a provenance field

harness 的自演化循环与本文同构:任务产生数据、validators 产生判定、判定又反过来塑造后续任务。可借鉴的有三点。第一,给每条准入样本与每条校验记录加上来源标识(源仓库、源任务族、源日期),让「这个校验器有没有见过这个来源」可以被程序化判断——没有这个字段,源级交叉拟合无法实施。第二,把「一致率」拆开:现在 verification_eval.py 测的是漏检率与误杀率,建议再加一个「同一错误上的一致率」,也就是闸门与生成器一起判错的比例,这正是 F 在 harness 语境下的对应物。第三,审计要与循环解耦:用一批冻结样本离线复核闸门,而不是让闸门自己判自己。

The harness self-evolution loop is isomorphic to the one here: tasks produce data, validators produce verdicts, and verdicts reshape later tasks. Three things transfer. First, attach a source identifier (repo, task family, date) to every admitted sample and every validation record so that whether a validator has seen a source becomes a programmatic question; without that field, source-level cross-fitting cannot be implemented at all. Second, split the agreement rate: verification_eval.py currently measures misses and false kills, so add agreement on the same error, the share of cases where the gate and the generator are wrong together, which is the harness analogue of F. Third, keep the audit decoupled: re-check the gate offline on a frozen sample instead of letting it judge itself.

Jev 备考:出题者与判分者是同一个闭环

Jev Exam Prep: the question writer and the grader form one loop

这个项目里,模型既出题(生成题干与 rubric 得分点)又判分,是本文闭环的教科书实例:如果某个得分点的参考答案来自对材料的一次系统性误读,它既会写进 rubric,又会在判分时被同一套理解认下来——出题与判分互相印证,但两边都错。可落地的做法是给每个得分点保留它所属的材料片段标识,评测时区分「同一材料派生的题」与「其他材料的题」,比较两组的判分准确率;如果同源组显著更低,就说明存在本文意义上的假一致。这与上一篇 JEV-as-a-Judge 里「文风对抗会击穿置信度」的结论互补:那篇说评测者会失灵,这篇说怎么发现并切断失灵的回流。

In that project one model both writes the questions (stems and rubric points) and grades the answers, making it a textbook instance of this loop: if a rubric point reference answer comes from a systematic misreading of the material, that misreading is written into the rubric and then confirmed at grading time, so writer and grader corroborate each other while both are wrong. Concretely, keep the material-span identifier for every rubric point, then compare grading accuracy between questions derived from the same material and questions from other material; a materially lower same-source accuracy indicates false agreement in this paper sense. This complements the earlier JEV-as-a-Judge finding that adversarial style breaks confidence: that paper shows the evaluator fails, this one shows how to detect and cut the return path that hides the failure.

两个项目共用一条结论:自生成的训练信号必须按血缘审计——同源不一定同错,但同错往往同源。

Both projects share one conclusion: self-generated training signals have to be audited by provenance — the same source does not guarantee the same error, but the same error usually shares a source.

这篇论文的价值在于把「自演化闭环会自我欺骗」变成可测量的东西:它没有停在「奖励可能被黑」的一般论述,而是隔离出一条具体回流路径,给出可区分真一致与假一致的指标,并用机制消融与固定回放把归因做干净。

The value here is turning self-deception in self-evolution into something measurable. It does not stop at the general claim that rewards can be hacked; it isolates one concrete return path, defines metrics that separate real from false agreement, and uses mechanism ablations plus fixed replay to keep the attribution clean.

一句话:只检查标签不够,关键是评测者有没有见过这个来源;而且要小心——把这条回流路径修好之后,内部一致率会下降,那是正常的。 In one line: checking labels is not enough, what matters is whether the evaluator has seen that source; and be careful, because fixing the return path makes internal agreement go down, and that is expected.