用图文符号对齐方法,让通用模型轻松学会星系形态分析。
GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology Analysis
- 通过图像、符号、文本三模态对齐提升模型理解力。
- 在分类与检索任务中表现优于传统微调方法。
- 适合天文研究者快速构建高效星系分析工具。
星系形态分析旨在基于星系形状与结构进行研究,核心任务包括识别与分类天文图像中的星系,以及通过视觉或结构相似性检索星系。现有方法或在大规模标注数据上直接训练领域专用基础模型,或在小规模图像集上微调视觉基础模型。前者有效但成本高,后者资源效率高但准确率常偏低。为此,本文提出 GalaxAlign,一种受公民科学家通过文本描述和图示符号识别星系启发的多模态方法。该方法采用三模态对齐框架,在微调阶段对齐三类数据:(1)表示星系形态的图示符号,(2)对应符号的文本标签,(3)星系图像。通过引入多模态指令,GalaxAlign无需昂贵预训练即可显著提升微调效果。在星系分类与相似性搜索任务上的实验表明,该方法能有效利用领域特定的多模态知识,微调通用预训练模型以完成天文学任务。代码已开源:https://github.com/RapidsAtHKUST/GalaxAlign。
原文摘要 · Abstract (English)
Galaxy morphology analysis involves studying galaxies based on their shapes and structures. For such studies, fundamental tasks include identifying and classifying galaxies in astronomical images, as well as retrieving visually or structurally similar galaxies through similarity search. Existing methods either directly train domain-specific foundation models on large, annotated datasets or fine-tune vision foundation models on a smaller set of images. The former is effective but costly, while the latter is more resource-efficient but often yields lower accuracy. To address these challenges, we introduce GalaxAlign, a multimodal approach inspired by how citizen scientists identify galaxies in astronomical images by following textual descriptions and matching schematic symbols. Specifically, GalaxAlign employs a tri-modal alignment framework to align three types of data during fine-tuning: (1) schematic symbols representing galaxy shapes and structures, (2) textual labels for these symbols, and (3) galaxy images. By incorporating multimodal instructions, GalaxAlign eliminates the need for expensive pretraining and enhances the effectiveness of fine-tuning. Experiments on galaxy classification and similarity search demonstrate that our method effectively fine-tunes general pre-trained models for astronomical tasks by incorporating domain-specific multi-modal knowledge. Code is available at https://github.com/RapidsAtHKUST/GalaxAlign.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。