arXiv:2508.07251cs.CV2025-08AAAI被引 7

构建首个面向动态4D场景的问答基准,支持细粒度时空推理。

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

  • 提出统一的时空建模框架,融合动态与静态信息处理4D点云。
  • 在92.7万条带思维链的问答上实现领先性能,验证多模态时序建模有效。
  • 适合研究具身智能、自动驾驶中人机交互与轨迹预测的学者。

从第一人称视角理解动态4D场景——即随时间变化的三维空间结构——对人机交互、自主导航和具身智能至关重要。现有第一人称数据集虽包含动态场景,但缺乏统一的4D标注与面向任务的评估协议,尤其在物体与人体运动及其交互的细粒度时空推理方面存在空白。为此,我们提出EgoDynamic4D,一个新型问答基准,包含RGB-D视频、相机位姿、全局唯一实例掩码及4D边界框。构建了92.7万条问答对,并附带显式思维链(Chain-of-Thought),支持可验证的逐步推理。设计12个动态问答任务,涵盖代理运动、人-物交互、轨迹预测、关系理解与时序因果推理,采用细粒度多维评估指标。为应对这些任务,我们提出端到端时空推理框架,统一动态与静态场景信息,通过实例感知特征编码、时间和相机编码以及空间自适应下采样,将大规模4D场景压缩为适合大语言模型处理的标记序列。在EgoDynamic4D上的实验表明,该方法持续优于基线,验证了多模态时序建模在第一人称动态场景理解中的有效性。

原文摘要 · Abstract (English)

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets contain dynamic scenes, they lack unified 4D annotations and task-driven evaluation protocols for fine-grained spatio-temporal reasoning, especially on motion of objects and human, together with their interactions. To address this gap, we introduce EgoDynamic4D, a novel QA benchmark on highly dynamic scenes, comprising RGB-D video, camera poses, globally unique instance masks, and 4D bounding boxes. We construct 927K QA pairs accompanied by explicit Chain-of-Thought (CoT), enabling verifiable, step-by-step spatio-temporal reasoning. We design 12 dynamic QA tasks covering agent motion, human-object interaction, trajectory prediction, relation understanding, and temporal-causal reasoning, with fine-grained, multidimensional metrics. To tackle these tasks, we propose an end-to-end spatio-temporal reasoning framework that unifies dynamic and static scene information, using instance-aware feature encoding, time and camera encoding, and spatially adaptive down-sampling to compress large 4D scenes into token sequences manageable by LLMs. Experiments on EgoDynamic4D show that our method consistently outperforms baselines, validating the effectiveness of multimodal temporal modeling for egocentric dynamic scene understanding.

4D点云时空推理具身智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。