让视觉语言模型理解宇宙尺度,突破传统几何限制。
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
- 设计几何提示与适配器,融合球面与双曲空间建模
- 在星系属性估计中达0.91的R²,在分类上提升0.17 F1
- 适合天体物理与跨尺度多几何场景研究者
现代视觉语言模型(VLMs)基于欧几里得向量空间构建图像块嵌入与卷积主干网络。当将VLM扩展至星系尺度以理解天文现象时,行星轨道的球面空间与黑洞的双曲空间引入双重挑战:(a) 当前预训练模型局限于欧几里得空间,缺乏综合几何嵌入;(b) 主流架构缺乏对各向异性物理几何的合适主干。本文提出银河漫步者(Galaxy-Walker),一种面向宇宙级视觉理解任务的几何感知VLM。我们设计几何提示,通过多尺度物理图上的随机游走生成几何标记,并引入几何适配器,以专家混合方式压缩并重塑空间各向异性。大量实验表明该方法有效,银河漫步者在星系属性估计任务中达到最高0.91的R²得分,在形态分类任务中对困难特征实现最高+0.17的F1提升,显著优于领域专用模型与通用VLM。
原文摘要 · Abstract (English)
Modern vision-language models (VLMs) develop patch embedding and convolution backbone within vector space, especially Euclidean ones, at the very founding. When expanding VLMs to a galaxy scale for understanding astronomical phenomena, the integration of spherical space for planetary orbits and hyperbolic spaces for black holes raises two formidable challenges. a) The current pre-training model is confined to Euclidean space rather than a comprehensive geometric embedding. b) The predominant architecture lacks suitable backbones for anisotropic physical geometries. In this paper, we introduced Galaxy-Walker, a geometry-aware VLM, for the universe-level vision understanding tasks. We proposed the geometry prompt that generates geometry tokens by random walks across diverse spaces on a multi-scale physical graph, along with a geometry adapter that compresses and reshapes the space anisotropy in a mixture-of-experts manner. Extensive experiments demonstrate the effectiveness of our approach, with Galaxy-Walker achieving state-of-the-art performance in both galaxy property estimation ($R^2$ scores up to $0.91$) and morphology classification tasks (up to $+0.17$ F1 improvement in challenging features), significantly outperforming both domain-specific models and general-purpose VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。