评测视觉语言模型在天文观测中的表现,发现物理知识是关键。
A systematic evaluation of vision-language models for observational astronomical reasoning tasks

- 构建5类天文任务的4100+实例基准,测试多模态模型表现
- 物理提示比现象提示更有效,数值表格优于图像提升13个百分点
- 模型常误判但自洽,需物理推理才能可信用于科研
视觉语言模型(VLMs)被广泛视为通用科学数据解读工具,但在真实天文观测中跨模态的可靠性尚未验证。我们提出AstroVLBench,一个包含超过4,100个专家验证实例的综合性基准,涵盖光学成像、射电干涉、多波段测光、时域光变曲线和光学谱学五项任务。评估六种前沿模型发现,性能显著依赖模态:其中Gemini 3 Pro在多数任务中表现最稳定,但各模型均显著低于领域专用方法。机制分析显示,性能不仅取决于关注显著视觉特征,更依赖将这些特征与物理知识对齐。描述‘看什么’的现象提示可提升准确率,而解释‘为什么重要’的物理提示效果更优,且分类更均衡、偏差更小。直接以数值表格替代渲染图像,最高可提升13个百分点。推理质量分析表明,缺乏物理对齐时,模型可能通过表面合理线索达成正确预测,但解释不具物理准确性,说明仅靠准确率不足以支撑可信科学应用。本研究首次系统建立多模态天文观测下VLMs的基准,揭示当前模型在表示、对齐与推理上的瓶颈。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astronomical observations across diverse modalities remains untested. We present AstroVLBench, a comprehensive benchmark comprising over 4,100 expert-verified instances across five tasks spanning optical imaging, radio interferometry, multi-wavelength photometry, time-domain light curves, and optical spectroscopy. Evaluating six frontier models, we find that performance is strongly modality-dependent: while one model (Gemini 3 Pro) emerges as the most consistently capable across tasks, task-specific strengths vary, and all models substantially underperform domain-specialized methods. Mechanistic ablations reveal that performance depends not only on directing attention to salient visual features but also on grounding those features in physical knowledge. Phenomenological prompts describing what to look for improve accuracy by sharpening model focus, but physical prompts explaining why those features matter perform better overall and yield more balanced classifications with reduced class-specific bias. Consistent with this picture, presenting the underlying one-dimensional measurements directly as numerical tables instead of rendered plots yields up to 13 percentage points improvement. Reasoning quality analysis further demonstrates that, without explicit physical grounding, models may reach correct predictions from phenomenologically plausible cues while providing physically imprecise justifications, establishing that accuracy alone is insufficient for trustworthy scientific deployment. These findings provide the first systematic, multi-modal baselines for VLMs in observational astronomy and identify the specific representation, grounding, and reasoning bottlenecks where current models fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。