arXiv:2608.30279cs.CV2026-08

通过互补掩码建模,提升点云视频的自监督表征学习效果

Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding

论文配图:Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding
图 1 · 摘自论文原文
  • 基于运动显著性分阶段掩码,聚焦关键动态区域
  • 在so(3)李代数中显式建模局部刚性旋转,增强运动感知
  • 跨视图令牌一致性预测,提升表征鲁棒性,适合动态3D场景理解

点云视频表征学习对三维动态场景理解至关重要。本文提出MoSaiC,一种新型运动-显著性互补掩码建模框架,用于自监督点云视频表征学习。MoSaiC包含三个组件:课程化运动-显著性掩码(CMSM),在课程调度下引导掩码过程聚焦于运动显著性标记;法向流运动(NFM)建模,在李代数so(3)中以显式几何运动目标监督每个标记的局部刚性旋转;跨视图令牌一致性预测(CTCP),在标记级别强制两个互补掩码视图间的一致性。三者协同使MoSaiC有效捕捉外观与运动动态。在动作识别、时间动作分割和点级语义分割等下游任务上的大量实验验证了方法的有效性。

原文摘要 · Abstract (English)

Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud video representation learning. MoSaiC couples three components: Curriculum Motion-Saliency Masking (CMSM), which guides the masking process toward motion-salient tokens under a curriculum schedule; Normal-Flow Motion (NFM) modeling, which supervises the local rigid rotation of each token in the Lie algebra so(3) as an explicit geometric motion target; and Cross-view Token Consistency Prediction (CTCP), which enforces consistency between two complementary masked views at the token level. Together, these components allow MoSaiC to effectively capture both appearance and motion dynamics. Extensive experiments on multiple downstream tasks, including action recognition, temporal action segmentation, and point-level semantic segmentation, demonstrate the effectiveness of our approach.

点云视频自监督学习运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。