AI在天体物理工作流中常无声生成错误结果,看似合理实则不准。
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows

- 在单次任务中,使用领域上下文使性能提升约6倍(0.85比~0)
- 压力测试下频繁出现无声失败,生成物理解释不一致的后验结果
- 适合关注科学AI可靠性与错误检测的研究者阅读
Agentic AI系统正越来越多地融入科学工作流,但其在真实条件下的表现仍不清晰。我们评估了CMBAgent在两种工作流范式和十八个天体物理任务上的表现。在单次任务设置中,使用领域特定上下文带来约6倍的性能提升(0.85对比~0),主要失败模式为沉默错误计算——语法正确但结果看似合理实则错误。在深度研究设置中,系统在压力测试中频繁出现沉默故障,生成物理不一致的后验分布且无自我诊断能力。总体而言,系统在定义明确的任务上表现良好,但在探测推理极限的问题上性能下降,且无明显错误信号。这些发现表明,最危险的失败模式并非显式错误,而是自信地生成错误结果。我们发布了评估框架,以促进科学AI代理的系统性可靠性分析。
原文摘要 · Abstract (English)
Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently understood. We evaluate CMBAgent across two workflow paradigms and eighteen astrophysical tasks. In the One-Shot setting, access to domain-specific context yields an approximately ~6x performance improvement (0.85 vs. ~0 without context), with the primary failure mode being silent incorrect computation - syntactically valid code that produces plausible but inaccurate results. In the Deep Research setting, the system frequently exhibits silent failures across stress tests, producing physically inconsistent posteriors without self-diagnosis. Overall, performance is strong on well-specified tasks but degrades on problems designed to probe reasoning limits, often without visible error signals. These findings highlight that the most concerning failure mode in agentic scientific workflows is not overt failure, but confident generation of incorrect results. We release our evaluation framework to facilitate systematic reliability analysis of scientific AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。