用三维空间分解提升2D图像的3D理解能力
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

- 将3D空间拆分为位置、方向、几何三类潜在表示
- 在复杂场景中实现比当前最优方法更高的推理准确率
- 适合需要精细空间推理的视觉语言模型研究者
尽管多模态大语言模型取得了显著进展,但从二维图像理解三维空间关系仍是关键挑战。现有方法主要依赖符号化文本标记,难以精确表达连续几何信息。虽然近期方法使用潜在表示增强推理,但单一潜在类型无法适应多样化的空间任务,在复杂几何场景中易出现错位。为此,我们提出GeoAnchor,一种交错式文本-潜在推理框架。该框架将三维空间信息分解为三个互补组件:用于物体定位的位置潜在表示、用于关系方向的方向潜在表示、以及用于场景结构的几何潜在表示。这些组件在结构化空间中重新组合,构建局部证据并捕捉全局上下文,实现动态且可解释的推理。此外,我们设计了一种协同训练策略,引导模型从局部空间感知逐步迈向全面的三维理解。在多种复杂三维推理任务上的实验表明,GeoAnchor超越了当前最优方法,验证了其有效性与泛化能力。
原文摘要 · Abstract (English)
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing methods primarily rely on symbolic text tokens, which inherently lack the fidelity to represent continuous geometric information. While recent methods use latent representations to enhance reasoning, relying on a single latent type cannot adapt to the diversity of spatial tasks, leading to misalignment in complex geometric scenarios. To address these limitations, we propose GeoAnchor, an interleaved text-latent reasoning framework. GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure. These components are recombined in a structured space to construct local evidence while capturing global context, enabling dynamic and interpretable reasoning. Furthermore, we introduce a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding. Extensive experiments on diverse and complex 3D reasoning tasks demonstrate that GeoAnchor outperforms the state of the art, validating its effectiveness and generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。