AI科学家答案对但理由错,需警惕机制错误导致的误导
Position: Correct Answer, Wrong Mechanism -- When AI Scientists Defend General Claims Their Own Data Contradicts
- 用28次实验检验代码代理在粒子识别中的推理过程
- 4/20主模型和3/8跨模型案例出现正确答案但错误机制
- 提出轻量级检测方法,可识别过度泛化的错误推理
当前将AI科学系统视为工具或合作者,仅关注最终结果,但本文指出这种仅看结果的评估方式不足。研究基于28个编码代理在Geant4模拟中重发现已知粒子识别可观测性的任务,包括对两个前沿模型的8次跨模型探测。结果显示,在20次主模型实验中有4次、8次跨模型实验中有3次,代理虽得出看似正确的结果,却依赖错误推理——条件变化后即失效,称为“正确答案,错误机制”(CAWM)。同一代理轨迹中,诚实性与机制一致性可分离:尽管所有五名代理在数据支持下拒绝了部分误导性先验,仍有一例以与自身数据矛盾的物理逻辑为所选可观测性辩护。在该仿真发现场景中,编码代理是可靠的工具,但不可靠的科学合作者,尤其在开放性主张生成任务中。可信合作需验证机制一致性,而现有代理无法自证。该失败可检测,本文提出轻量级测试:一步制式转变检查仅需代理主张即可标记过度泛化情况;若已知正确可观测性,则辅助重计算可覆盖剩余案例。二者联合可标记本研究中所有CAWM实例。
原文摘要 · Abstract (English)
AI scientist systems are described as tools, coauthors, or founders, but we evaluate them as if only the final answer matters. This position paper argues that outcome-only evaluation is insufficient, and that task outcome, mechanism fidelity, and epistemic honesty must be measured separately. Our evidence comes from 28 episodes of a coding agent attempting to rediscover a known particle identification observable in a Geant4 simulation, including an 8-episode probe across two additional frontier models. In 4/20 primary-model and 3/8 cross-model episodes, agents reach right-looking results through incorrect reasoning that breaks when conditions change, which we call Correct Answer, Wrong Mechanism (CAWM). Honesty and mechanism fidelity dissociate within a single agent trajectory. When given a partially misleading prior, all five agents reject the false component on evidence, yet one defends its chosen observable with physics inconsistent with its own data. In the simulation-based discovery setting studied here, coding agents prove reliable tools but unreliable scientific co-authors for open-ended claim-making, where co-author trust requires mechanism-fidelity verification they do not reliably self-apply. The failure is detectable, and we propose a lightweight test. A one-step regime-shift check needs only the agent's claim and flags the over-generalized cases. A companion recomputation flags the remaining cases when the correct observable is known. Together, these checks flag every CAWM case in this study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。