arXiv:2604.12630cs.CVcs.CL2026-04被引 2

通过动态整合多层几何特征,提升大模型空间推理能力

GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning

论文配图:GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
图 1 · 摘自论文原文
  • 构建分层几何特征库,用视觉令牌作为查询实现逐层稀疏路由
  • 在VSI-Bench等三个基准上超越更大模型,40亿参数达顶尖水平
  • 适合需要精准空间理解的视觉问答与3D场景分析任务

多模态大语言模型(MLLM)在各类视觉任务中表现卓越,但在空间推理方面仍存短板。现有方法通过引入3D基础模型的几何特征缓解问题,但依赖静态单层提取。我们发现该方式导致任务错位偏差:几何特征自然向3D预训练目标演化,可能违背MLLM多样化的空间需求,使单一层次始终不足。为此,我们提出GeoAlign框架,动态聚合多层几何特征以对齐实际需求。该方法构建分层几何特征库,利用MLLM原始视觉标记作为内容感知查询,实现逐层稀疏路由,自适应获取每块图像对应的合适几何特征。在VSI-Bench、ScanQA和SQA3D上的大量实验表明,我们的小型40亿参数模型有效达到当前最优性能,甚至超越更大规模的现有MLLM。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foundation models, but rely on static single-layer extractions. We identify that such an approach induces a task misalignment bias: the geometric features naturally evolve towards 3D pretraining objectives, which may contradict the heterogeneous spatial demands of MLLMs, rendering any single layer fundamentally insufficient. To resolve this, we propose GeoAlign, a novel framework that dynamically aggregates multi-layer geometric features to realign with the actual demands. GeoAlign constructs a hierarchical geometric feature bank and leverages the MLLM's original visual tokens as content-aware queries to perform layer-wise sparse routing, adaptively fetching the suitable geometric features for each patch. Extensive experiments on VSI-Bench, ScanQA, and SQA3D demonstrate that our compact 4B model effectively achieves state-of-the-art performance, even outperforming larger existing MLLMs.

空间推理几何特征多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。