无需显式运动信息,统一学习点云视频的几何与动态表征
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
- 通过隐式语义对齐学习运动,不依赖外部标注
- 自解耦策略分离低层几何与高层语义,提升表征质量
- 无需微调即可在动作分割等任务中达到领先效果
点云视频的自监督表征学习面临两大挑战:(1) 现有方法依赖显式运动知识,导致表征不优;(2) 传统掩码自编码器(MAE)难以衔接4D数据中低层几何与高层动态。本文提出一种新型自解耦MAE框架,通过隐式对齐潜在空间中的高层语义来学习运动,无需任何显式知识。引入自解耦学习策略,将几何令牌与潜在令牌共同输入共享解码器,有效分离低层几何与高层语义。除重建目标外,还设计帧级运动与视频级全局信息三类对齐目标以增强时序理解。实验表明,预训练编码器无需微调即可出色区分时空表征。在MSR-Action3D、NTU-RGBD、HOI4D、NvGesture和SHREC'17上广泛验证,无论粗粒度还是细粒度任务均表现优异,尤其在HOI4D动作分割任务上提升3.8%准确率。
原文摘要 · Abstract (English)
Self-supervised representation learning for point cloud videos remains a challenging problem with two key limitations: (1) existing methods rely on explicit knowledge to learn motion, resulting in suboptimal representations; (2) prior Masked AutoEncoder (MAE) frameworks struggle to bridge the gap between low-level geometry and high-level dynamics in 4D data. In this work, we propose a novel self-disentangled MAE for learning expressive, discriminative, and transferable 4D representations. To overcome the first limitation, we learn motion by aligning high-level semantics in the latent space \textit{without any explicit knowledge}. To tackle the second, we introduce a \textit{self-disentangled learning} strategy that incorporates the latent token with the geometry token within a shared decoder, effectively disentangling low-level geometry and high-level semantics. In addition to the reconstruction objective, we employ three alignment objectives to enhance temporal understanding, including frame-level motion and video-level global information. We show that our pre-trained encoder surprisingly discriminates spatio-temporal representation without further fine-tuning. Extensive experiments on MSR-Action3D, NTU-RGBD, HOI4D, NvGesture, and SHREC'17 demonstrate the superiority of our approach in both coarse-grained and fine-grained 4D downstream tasks. Notably, Uni4D improves action segmentation accuracy on HOI4D by $+3.8\%$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。