arXiv:2509.02615astro-ph.IMcs.AI2025-09被引 5

用通用视觉语言模型识别射电星系形态,发现提示词敏感但微调后效果接近专业模型。

Radio Astronomy in the Era of Vision-Language Models: Prompt Sensitivity and Adaptation

  • 用自然语言和图示设计提示词,测试模型对射电图像的分类能力
  • 微调后仅1500万参数即达3%错误率,接近顶尖专业模型
  • 提示词微小变化导致结果剧烈波动,警示科学应用中的可靠性风险

视觉语言模型(如Qwen、Gemini)被视为跨领域通用人工智能系统,但在科学成像中面对陌生数据分布时的能力仍不明确。本文评估通用VLM在未接触天文数据的前提下,能否基于形态对射电星系进行分类,使用MiraBest FR-I/FR-II数据集。研究探索了自然语言与示意图提示策略,并首次在天文学中引入视觉上下文样例作为提示。同时,采用轻量级监督微调(LoRA)进行适应。结果显示:(i) 即使仅靠提示也取得良好性能,表明VLM对陌生科学领域具备有用先验;(ii) 但输出极不稳定,仅改变布局、顺序或解码温度等表面因素,结果便剧烈波动,即便语义不变;(iii) 仅需1500万可训练参数且无天文预训练,微调后的Qwen-VL即可达到3%错误率,接近现有最优水平。这表明,当前所谓“推理”常是提示敏感性所致,而非真实推断。尽管如此,经极简适配后,通用模型已能媲美专用模型,为科学发现提供有潜力但脆弱的新工具。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs), such as recent Qwen and Gemini models, are positioned as general-purpose AI systems capable of reasoning across domains. Yet their capabilities in scientific imaging, especially on unfamiliar and potentially previously unseen data distributions, remain poorly understood. In this work, we assess whether generic VLMs, presumed to lack exposure to astronomical corpora, can perform morphology-based classification of radio galaxies using the MiraBest FR-I/FR-II dataset. We explore prompting strategies using natural language and schematic diagrams, and, to the best of our knowledge, we are the first to introduce visual in-context examples within prompts in astronomy. Additionally, we evaluate lightweight supervised adaptation via LoRA fine-tuning. Our findings reveal three trends: (i) even prompt-based approaches can achieve good performance, suggesting that VLMs encode useful priors for unfamiliar scientific domains; (ii) however, outputs are highly unstable, i.e. varying sharply with superficial prompt changes such as layout, ordering, or decoding temperature, even when semantic content is held constant; and (iii) with just 15M trainable parameters and no astronomy-specific pretraining, fine-tuned Qwen-VL achieves near state-of-the-art performance (3% Error rate), rivaling domain-specific models. These results suggest that the apparent "reasoning" of VLMs often reflects prompt sensitivity rather than genuine inference, raising caution for their use in scientific domains. At the same time, with minimal adaptation, generic VLMs can rival specialized models, offering a promising but fragile tool for scientific discovery.

视觉语言模型射电天文学提示工程微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。