用点云、骨骼、惯性数据和文本联合建模,提升点云人体动作理解能力
DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding
- 构建四模态联合嵌入空间,融合点云、骨骼、IMU与文本
- 在LIPD+Babel数据集上实现跨模态匹配与时序定位,准确率超基线
- 适用于隐私保护场景下的动作识别与检索,适合多模态感知研究者
尽管激光雷达(LiDAR)是替代RGB相机进行人体动作感知的隐私友好方案,但在多模态对比预训练方面仍研究不足。本文提出DeSPITE模型,通过联合学习点云、人体骨骼姿态、惯性测量单元(IMU)数据与文本的对应关系,在统一嵌入空间中实现跨模态对齐。我们整合了LIPD与Babel数据集,实现四模态数据同步,首次支持在点云序列上开展骨骼<->点云<->IMU匹配、动作检索与时序片段检索等新任务。实验表明,DeSPITE在MSR-Action3D与HMPEAR数据集上显著提升点云动作识别性能,是一种有效的点云动作理解预训练策略。
原文摘要 · Abstract (English)
Despite LiDAR (Light Detection and Ranging) being an effective privacy-preserving alternative to RGB cameras to perceive human activities, it remains largely underexplored in the context of multi-modal contrastive pre-training for human activity understanding (e.g., human activity recognition (HAR), retrieval, or person re-identification (RE-ID)). To close this gap, our work explores learning the correspondence between LiDAR point clouds, human skeleton poses, IMU data, and text in a joint embedding space. More specifically, we present DeSPITE, a Deep Skeleton-Pointcloud-IMU-Text Embedding model, which effectively learns a joint embedding space across these four modalities. At the heart of our empirical exploration, we have combined the existing LIPD and Babel datasets, which enabled us to synchronize data of all four modalities, allowing us to explore the learning of a new joint embedding space. Our experiments demonstrate novel human activity understanding tasks for point cloud sequences enabled through DeSPITE, including Skeleton<->Pointcloud<->IMU matching, retrieval, and temporal moment retrieval. Furthermore, we show that DeSPITE is an effective pre-training strategy for point cloud HAR through experiments in MSR-Action3D and HMPEAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。