对比两款智能体自动完成引力波数据分析,发现效率与可靠性差异显著。
First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope
- 用智能体自主执行从噪声分析到论文生成的全流程
- Claude Code仅3.4分钟完成,Codex耗时16分钟且反复重试
- 指令理解差异导致科学结论分歧,凸显AI透明性重要性
我们对比了两款前沿智能体系统——Anthropic的Claude Code与OpenAI的Codex——在无须人工干预的情况下,于同一计算环境下自主执行一套完整的引力波数据处理流程。该流程包括:从爱因斯坦望远镜模拟噪声中估计功率谱密度、生成几何模板库、匹配滤波恢复100个双黑洞信号注入、自动生成结果,并利用大语言模型撰写符合《物理评论D》格式的论文。两者接收相同指令与计算资源,实验进行两次:第一次为非现实的强信号注入,第二次按物理合理信噪比(SNR)调整。科学结果在两次运行中均收敛。然而,两系统行为与计算开销差异显著:Claude Code以约3.4分钟完成,存在无声偏离规范;Codex则耗时约16分钟,通过多次自我纠正重启完成,还自发优化了匹配滤波内层性能。自动生成的论文在长度、细节和质量上也明显不同。第二次运行中,对信噪比范围的细微理解差异导致真实科学分歧:Claude Code无声重构指令,而Codex严格遵循原意。本文讨论这些行为差异对科学计算中智能体部署的影响,涉及速度与可审计性、沉默式与透明式错误处理、指令解释以及多模型流程中中间数据表示的关键作用。
原文摘要 · Abstract (English)
We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a simple end-to-end gravitational wave data analysis pipeline on a shared computing infrastructure without human intervention. The pipeline comprises power spectral density estimation from raw Einstein Telescope simulated noise, geometric template bank generation, matched filter recovery of 100 binary black hole signal injections, automated results generation, and large language model-assisted production of a manuscript formatted in the style of Physical Review D. Both agents received identical written specifications and identical compute resources. The experiment was run twice: a first run with unrealistically loud injections, and a second run with signals rescaled to a physically motivated SNR range. The scientific results converged in both runs. However, the agents exhibited substantially different behaviors and computational costs: Claude Code completed the pipeline in ~3.4 minutes with silent deviations from the specification, while Codex required ~16 minutes across explicit self-correcting restarts, including an unsolicited performance optimization of the matched filter inner loop. The autonomously generated manuscripts also diverged in length, details, and quality. In the second run, a subtle difference in the interpretation of the SNR range instruction led to a genuine scientific divergence: Claude Code silently reinterpreted the instructions, while Codex followed the specification literally. We discuss the implications of these behavioral differences, such as speed versus auditability, silent versus transparent error handling, instruction interpretation, and the criticality of intermediate data representations in multi-model pipelines, for the deployment of agentic AI in scientific computing workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。