arXiv:2511.22961cs.CV2025-11

用多视图图像和带坐标描述的文本,提升大模型对3D场景的理解能力

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

  • 输入层显式对齐视觉与语言模型,融合多视角图像和带坐标的文本
  • 在两类3D问答任务上均超越现有方法,实现更精准的场景推理
  • 适合需要理解复杂空间关系的3D视觉任务,如机器人导航

大型视觉-语言模型(VLM)在3D场景理解方面展现出巨大潜力。现有方法通常将3D场景特征隐式对齐至VLM的嵌入空间,但受限于3D数据稀缺及空间关系复杂性,表现不佳。为此,本文提出一种新的分层多模态表示方法,通过融合多视图图像与文本描述,在输入层显式对齐至VLM。文本描述通过引用检测到物体的3D坐标来捕捉空间关系;多视图图像包含俯视图及前、左、右、后四个方向视图,确保场景全覆盖。同时,引入分层特征表示,将局部图像块特征聚合为视图级和场景级表示,支持局部与全局上下文联合推理。在情境化3D问答与通用3D问答基准上的实验结果验证了该方法的有效性。

原文摘要 · Abstract (English)

Recent advances in large vision-language models (VLMs) have shown significant promise for 3D scene understanding. Existing VLM-based approaches typically align 3D scene features with the VLM's embedding space. However, this implicit alignment often yields suboptimal performance due to the scarcity of 3D data and the inherent complexity of spatial relationships in 3D environments. To address these limitations, we propose a novel hierarchical multimodal representation for 3D scene reasoning that explicitly aligns with VLMs at the input space by leveraging both multi-view images and text descriptions. The text descriptions capture spatial relationships by referencing the 3D coordinates of detected objects, while the multi-view images include a top-down perspective and four directional views (forward, left, right, and backward), ensuring comprehensive scene coverage. Additionally, we introduce a hierarchical feature representation that aggregates patch-level image features into view-level and scene-level representations, enabling the model to reason over both local and global scene context. Experimental results on both situated 3D Q&A and general 3D Q&A benchmarks demonstrate the effectiveness of our approach.

3D理解多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。