自动化研究中,一次实验结果可能误导对想法的判断。
One Run Is Not an Idea: The Implementation Lottery in Automated Research

- 用多次实现验证想法可靠性,避免单一实验偏差。
- 同一任务下,不同实现差异是重复实验误差的5~10倍。
- 适合评估自动化研究系统可信度的研究者或团队。
自动化研究系统用实验得分决定产出物和保留想法。但一次运行仅对应一个想法的实现版本。将此实现层面的得分视为对核心机制的证据,会引发‘实现抽奖’问题——想法结论依赖于偶然采样的实现方式。这种不匹配在每次运行更新机制信念时均存在。我们提出‘想法可靠性审计’:冻结候选卡片,用无结果感知的保真度标签,重新采样会话级实现并重跑保存产物,报告想法ICC与留一实现排除(LOO)胜者反转率。在13个表格任务和两个编码代理设置的312次任务中,实现方差分别超过同产物重跑方差的5倍和10倍,且单次实现抽样胜者与另两次平均胜者不同的决策占比达25.6%和43.6%。两种无结果感知评审规则下,反转仍存。对三个材料回归工作流的诊断也显示,实现差异主导了分解结果。这些发现区分了想法可靠性与最佳N项产物性能。在分数用于指导想法分支、迁移或研究记忆前,需覆盖多实现证据。
原文摘要 · Abstract (English)
Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。