首个面向动态点云理解的多模态大模型,支持时序推理与动作分析。
4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
- 基于双阶段构建4D点云序列与问答对,实现动态场景建模。
- 在20万组问答上验证,动作理解与时序推理性能显著提升。
- 引入故障感知自举机制,可自动发现并修复模型短板。
点云为3D物体提供了紧凑且丰富的表达,近年来被整合进多模态大语言模型(MLLM)。然而现有方法主要关注静态对象,动态点云序列的理解仍鲜有探索。这一局限主要源于缺乏大规模跨模态数据集,以及在时空上下文中建模运动的困难。为此,我们提出4DPC$^2$hat,首个专用于动态点云理解的MLLM。我们通过精心设计的两阶段流程——拓扑一致的4D点云构建与两级标注,构建了大规模跨模态数据集4DPC$^2$hat-200K。该数据集包含超过44,000个动态物体序列、70万帧点云及20万条精选问答对,支持计数、时间关系、动作、空间关系和外观等方面的提问。核心框架中,我们引入增强型Mamba时序推理MLLM,以捕捉点云序列中的长程依赖与动态模式。此外,提出故障感知自举学习策略,迭代识别模型缺陷并生成针对性问答监督,持续强化推理能力。大量实验表明,相比现有模型,4DPC$^2$hat在动作理解与时序推理方面均有显著提升,为4D动态点云理解奠定坚实基础。
原文摘要 · Abstract (English)
Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing methods primarily focus on static objects, while understanding dynamic point cloud sequences remains largely unexplored. This limitation is mainly caused by the lack of large-scale cross-modal datasets and the difficulty of modeling motions in spatio-temporal contexts. To bridge this gap, we present 4DPC$^2$hat, the first MLLM tailored for dynamic point cloud understanding. To this end, we construct a large-scale cross-modal dataset 4DPC$^2$hat-200K via a meticulous two-stage pipeline consisting of topology-consistent 4D point construction and two-level captioning. The dataset contains over 44K dynamic object sequences, 700K point cloud frames, and 200K curated question-answer (QA) pairs, supporting inquiries about counting, temporal relationship, action, spatial relationship, and appearance. At the core of the framework, we introduce a Mamba-enhanced temporal reasoning MLLM to capture long-range dependencies and dynamic patterns among a point cloud sequence. Furthermore, we propose a failure-aware bootstrapping learning strategy that iteratively identifies model deficiencies and generates targeted QA supervision to continuously strengthen corresponding reasoning capabilities. Extensive experiments demonstrate that our 4DPC$^2$hat significantly improves action understanding and temporal reasoning compared with existing models, establishing a strong foundation for 4D dynamic point cloud understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。