arXiv:2510.11509cs.CV2025-10NeurIPS被引 3

构建首个情境化3D变化理解数据集,助力多模态大模型理解动态环境。

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

  • 基于1.1万次人类观察构建共享情境认知,融合第一/第三人称视角与空间关系。
  • 包含12.1万问答对、3.6万变化描述和1.7万重排指令,覆盖感知-行动全链路任务。
  • 提出轻量级3D多模态模型SCReasoner,高效对比点云变化且无额外参数开销。

物理环境本质上是动态的,但现有3D数据集和评估基准往往只关注动态场景或情境的单一维度,导致理解不完整。为此,我们提出Situat3DChange,一个支持三种情境感知变化理解任务的大规模数据集,遵循感知-动作模型:包含12.1万问答对、3.6万变化描述(用于感知任务)和1.7万重排指令(用于动作任务)。为构建该数据集,Situat3DChange利用1.1万次人类对环境变化的观察,建立人机协作所需的共享心智模型与情境意识。这些观察融合了第一人称与第三人称视角,以及类别化与坐标式空间关系,并通过大语言模型整合,以支持情境化变化理解。针对同一场景中微小点云变化的对比难题,我们提出SCReasoner——一种高效的3D多模态大模型方法,可在极低参数开销下完成有效点云比较,且无需为语言解码器增加额外标记。在Situat3DChange上的全面评估揭示了多模态大模型在动态场景与情境理解中的进展与局限。额外实验表明,使用Situat3DChange作为训练数据集,在数据扩展与跨领域迁移方面均展现出任务无关的有效性。

原文摘要 · Abstract (English)

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange, an extensive dataset supporting three situation-aware change understanding tasks following the perception-action model: 121K question-answer pairs, 36K change descriptions for perception tasks, and 17K rearrangement instructions for the action task. To construct this large-scale dataset, Situat3DChange leverages 11K human observations of environmental changes to establish shared mental models and shared situational awareness for human-AI collaboration. These observations, enriched with egocentric and allocentric perspectives as well as categorical and coordinate spatial relations, are integrated using an LLM to support understanding of situated changes. To address the challenge of comparing pairs of point clouds from the same scene with minor changes, we propose SCReasoner, an efficient 3D MLLM approach that enables effective point cloud comparison with minimal parameter overhead and no additional tokens required for the language decoder. Comprehensive evaluation on Situat3DChange tasks highlights both the progress and limitations of MLLMs in dynamic scene and situation understanding. Additional experiments on data scaling and cross-domain transfer demonstrate the task-agnostic effectiveness of using Situat3DChange as a training dataset for MLLMs.

3D理解多模态情境感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。