arXiv:2601.06498cs.CLastro-ph.IM2026-01ACL

用AI助手自动分析星体光谱,提升罕见天体发现效率。

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

  • 结合视觉语言模型与天文工具,通过多模态推理模拟专家审阅流程。
  • 在五项罕见天体识别任务中,宏平均F1得分从28.3提升至76.5。
  • 可跨数据集泛化,适合天文研究者和自动化巡天项目使用。

由于深度学习分类器的泛化能力与可解释性有限,稀有天体候选者的最终确认仍依赖于天文学家的人工视觉检查——这一过程耗时且难以扩展。天文学家利用专业工具分析光谱以建立可靠星表,但面对现代光谱巡天产生的海量数据,该方法已成为主要瓶颈。为此,我们提出Spec-o3,一种工具增强型视觉语言代理,通过交错式多模态思维链推理实现与天文学家一致的光谱检查。Spec-o3采用两阶段后训练策略:先在专家检查轨迹上进行冷启动监督微调,再在稀有类型验证任务上进行基于结果的强化学习。在LAMOST的五个稀有天体识别任务上评估,其宏平均F1得分从28.3提升至76.5,使用70亿参数基础模型,优于专有视觉语言模型与专用深度模型。关键的是,该代理在跨巡天迁移(从LAMOST到SDSS/DESI)中表现出强泛化能力。专家评估确认其推理路径连贯且物理一致,支持透明可信的决策。代码、数据与模型已公开于https://github.com/Maxwell-Jia/spec-o3。

原文摘要 · Abstract (English)

Due to the limited generalization and interpretability of deep learning classifiers, The final vetting of rare celestial object candidates still relies on expert visual inspection--a manually intensive process. In this process, astronomers leverage specialized tools to analyze spectra and construct reliable catalogs. However, this practice has become the primary bottleneck, as it is fundamentally incapable of scaling with the data deluge from modern spectroscopic surveys. To bridge this gap, we propose Spec-o3, a tool-augmented vision-language agent that performs astronomer-aligned spectral inspection via interleaved multimodal chain-of-thought reasoning. Spec-o3 is trained with a two-stage post-training recipe: cold-start supervised fine-tuning on expert inspection trajectories followed by outcome-based reinforcement learning on rare-type verification tasks. Evaluated on five rare-object identification tasks from LAMOST, Spec-o3 establishes a new State-of-the-Art, boosting the macro-F1 score from 28.3 to 76.5 with a 7B parameter base model and outperforming both proprietary VLMs and specialized deep models. Crucially, the agent demonstrates strong generalization to unseen inspection tasks across survey shifts (from LAMOST to SDSS/DESI). Expert evaluations confirm that its reasoning traces are coherent and physically consistent, supporting transparent and trustworthy decision-making. Code, data, and models are available at https://github.com/Maxwell-Jia/spec-o3.

天体发现视觉语言模型光谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。