用天体物理约束测试AI在动态模型拟合中的表现。
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints

- 构建可扩展的仿真环境,评估AI在径向速度数据上的迭代拟合能力。
- 8个前沿智能体虽拟合效果好,但常偏离真实物理参数。
- 适合研究科学推理、模型拟合与智能体训练的科研人员使用。
自主AI智能体的兴起表明,需要具备科学任务反馈的动态基准环境来评估其科研能力。我们提出Stargazer,一个基于径向速度时间序列数据的可扩展环境,用于评估智能体在动态、迭代的物理约束模型拟合任务中的表现。该环境包含120个任务,分为三个难度层级,涵盖20个真实历史案例,涉及从高信噪比单行星系统到复杂多行星配置的多种场景,需处理低信噪比数据分析。对8个前沿智能体的评估显示,尽管它们在统计拟合上表现良好,但往往无法恢复正确的物理系统参数,这一局限性即使在赋予基础技能后依然存在。增加测试时计算资源仅带来微小提升,过量的令牌消耗常反映重复失败循环而非有效探索。Stargazer为训练、评估、指导和扩展模型拟合策略提供了实践相关的机会,其仿真驱动的设计方法可推广至多个科学领域的模型拟合问题。源代码与项目网站分别位于https://github.com/AIPS-UofT/Stargazer 和 https://aips-uoft.github.io/Stargazer/。
原文摘要 · Abstract (English)
The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to evaluate the capabilities of these agents in research work. We introduce Stargazer, a scalable environment for evaluating AI agents on dynamic, iterative physics-grounded model-fitting tasks using inference on radial-velocity (RV) time series data. Stargazer comprises 120 tasks across three difficulty tiers, including 20 real archival cases, covering diverse scenarios ranging from high-SNR single-planet systems to complex multi-planetary configurations requiring involved low-SNR analysis. Our evaluation of eight frontier agents reveals a gap between numerical optimization and adherence to physical constraints: although agents often achieve a good statistical fit, they frequently fail to recover correct physical system parameters, a limitation that persists even when agents are equipped with vanilla skills. Furthermore, increasing test-time compute yields only marginal gains, with excessive token usage often reflecting recursive failure loops rather than meaningful exploration. Stargazer presents an opportunity to train, evaluate, scaffold, and scale strategies on a model-fitting problem of practical research relevance today. Our methodology to design a simulation-driven environment for AI agents presumably generalizes to many other model-fitting problems across scientific domains. Source code and the project website are available at https://github.com/AIPS-UofT/Stargazer and https://aips-uoft.github.io/Stargazer/, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。