发现科研代理因只看总分而误选错误候选,提出外部审计机制。
Search Discipline for Long-Horizon Research Agents
- 用外部审计取代单一总分判断,审查每个区域的细节表现。
- 在生态系统模型中,高分候选会破坏森林,低分反而保护生态。
- 适合关注模型可靠性、避免表面指标误导的研究者使用。
自研科研代理通常根据一个综合评分选择科学候选,但该评分常将异质区域、子集或人群的多样性结果聚合。我们发现,当科学有效性存在于未聚合的结构中时,综合评分可能将错误候选排在首位。尽管总体得分提升,其底层结构却发生反转,导致基于分数的决策采纳了悄然破坏模型的候选。这种失败不局限于特定领域,只要候选的有效性多维而验证者仅依赖单一聚合,就会出现。我们在生态系统演化的模型中演示了这一反转:最高分候选与略低分候选在全局得分上差异在噪声范围内,但前者使受保护的北方森林崩溃,后者则保留它们。二者区别在于各区域行为,而非总分。因此,不应将决策权交给生成候选的代理。优化得分的代理最不可能察觉评分失效,且一旦代理停止,提示已无后续机会。我们引入外部控制循环,对每个候选的非聚合行为进行审计,并在代理决策后介入。它可降级代理本会接受的候选,也可重启代理已宣布结束的运行。本文贡献在于发现该反转现象,以及一种基于可审查候选效应证据而非得分的搜索纪律协议。
原文摘要 · Abstract (English)
Autoresearch agents now propose, evaluate, and select scientific candidates against a metric, and that metric is usually an aggregate reduced over a heterogeneous space of regions, slices, or cohorts. We show that when scientific validity lives in that disaggregated structure, the aggregate can rank the wrong candidate first. The headline number improves while the structure underneath inverts, so a decision made on the number accepts a candidate that quietly breaks the model. The failure is not domain-specific. It appears wherever a candidate's validity is multi-dimensional but its verifier is a single reduction. We demonstrate the inversion on a fire-model task in the Ecosystem Demography model. The highest-scoring candidate and a slightly lower one are within noise of each other on global score, yet the top-scoring one collapses the protected boreal regions while the other preserves them. What separates them is the per-region behavior, not the headline number. This decision should not be left to the agent that produced the candidates. The agent optimizing the score is the last party likely to catch the score being wrong, and a prompt has no remaining turn once the agent has stopped. We move the decision to an external control loop that audits each candidate on its disaggregated behavior and acts after the agent has decided. It can demote a candidate the agent would have accepted, and it can reopen a run the agent had declared finished. Our contribution is the inversion finding itself, and a search-discipline protocol that decides on reviewable candidate-effect evidence instead of the score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。