arXiv:2603.27507cs.CV2026-03TPAMI被引 2

用上下文丰富的物体序列提升3D场景理解,让大模型更懂复杂环境。

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

论文配图:Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
图 1 · 摘自论文原文
  • 将3D场景转为带语义的物体序列,实现以物体为中心的交互。
  • 在5个主流3D视觉语言任务上达到顶尖性能,无需额外微调。
  • 仅需2D输入即可实现真实场景应用,避免昂贵的3D重建。

多模态大语言模型(MLLM)在3D场景理解方面展现出巨大潜力,但现有方法在细粒度物体定位和上下文推理上表现不足,限制了其对复杂3D环境的解读与交互能力。本文提出Chat-Scene++,一种将3D场景表示为富含上下文语义的物体序列的MLLM框架。通过将场景分解为带有标识符的物体表征,使大模型能在多样化的3D视觉语言任务中响应指令。该框架利用大规模预训练的3D场景级与2D图像级编码器提取具有上下文信息的物体特征,而非孤立的单物体特征。其灵活的物体中心设计支持基于实体的思维链(G-CoT)推理,可在多步推理中区分物体的类别与空间位置。无需添加特定任务头部或微调,Chat-Scene++在五个主要3D视觉语言基准(ScanRefer、Multi3DRefer、Scan2Cap、ScanQA、SQA3D)上均取得当前最优表现,验证了其在场景理解、物体定位与空间推理方面的有效性。此外,仅依赖2D输入即可实现真实场景应用,无需进行计算成本高昂的3D重建。

原文摘要 · Abstract (English)

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, limiting their ability to interpret and interact with complex 3D environments. In this paper, we present Chat-Scene++, an MLLM framework that represents 3D scenes as context-rich object sequences. By structuring scenes as sequences of objects with contextual semantics, Chat-Scene++ enables object-centric representation and interaction. It decomposes a 3D scene into object representations paired with identifier tokens, allowing LLMs to follow instructions across diverse 3D vision-language tasks. To capture inter-object relationships and global semantics, Chat-Scene++ extracts context-rich object features using large-scale pre-trained 3D scene-level and 2D image-level encoders, unlike the isolated per-object features in Chat-Scene. Its flexible object-centric design also supports grounded chain-of-thought (G-CoT) reasoning, enabling the model to distinguish objects at both category and spatial levels during multi-step inference. Without the need for additional task-specific heads or fine-tuning, Chat-Scene++ achieves state-of-the-art performance on five major 3D vision-language benchmarks: ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D. These results highlight its effectiveness in scene comprehension, object grounding, and spatial reasoning. Additionally, without reconstructing 3D worlds through computationally expensive processes, we demonstrate its applicability to real-world scenarios using only 2D inputs.

3D理解多模态大模型物体定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。