通过预测掩码时空点管,实现4D点云视频的自监督学习。
Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos

- 用隐空间中的点管预测替代原始坐标重建。
- 在动作和手势识别任务上显著提升少样本与跨数据集性能。
- 适合做4D点云视频的预训练,尤其对标注稀缺场景有效。
4D点云视频的自监督表示学习面临标注成本高、基于重建的预训练易过度关注低层几何细节的问题。本文提出一种类似JEPA的框架,通过从可见上下文特征中预测被掩码的时空点管的潜在表示,实现无标签时空点云的学习。不直接重建原始坐标,而是利用隐空间中的目标表示进行预测。为稳定隐表示,引入了草图各向同性高斯正则化,避免嵌入坍塌且无需显式重建目标。该方法旨在同时捕捉空间结构与时间动态,并使预训练目标与下游语义识别保持一致。在动作与手势识别基准上的实验表明,所学表示在微调、小样本学习和跨数据集迁移上均有提升。结果表明,基于隐空间预测的JEPA范式是4D点云视频预训练的有力替代方案。
原文摘要 · Abstract (English)
Self-supervised representation learning for 4D point cloud videos is challenging because annotations are costly and reconstruction-based pretraining can overemphasize low-level geometric details. We propose a JEPA-style framework that learns from unlabeled spatiotemporal point clouds through latent point-tube prediction. Instead of reconstructing raw coordinates, the model masks spatiotemporal regions and predicts their target representations from visible context representations in feature space. To stabilize latent prediction, we incorporate Sketched Isotropic Gaussian Regularization, which encourages non-collapsed embeddings without relying on explicit reconstruction targets. This formulation aims to capture both spatial structure and temporal dynamics while keeping the pretraining objective aligned with downstream semantic recognition. Experiments on action and gesture recognition benchmarks show that the learned representations improve downstream fine-tuning, limited-label learning, and cross-dataset transfer. These results suggest that JEPA-style latent prediction is a promising alternative to reconstruction-centered pretraining for 4D point cloud videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。