用对齐+适配策略,让3D模型高效学会看4D点云视频。
Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception
- 先用最优传输对齐3D与4D数据分布,再通过轻量适配增强时序建模。
- 在4D动作分割上提升8.7%,4D语义分割达84.06%,参数量远低于全微调。
- 适合资源有限但需处理动态点云视频的机器人与自动驾驶场景。
点云视频理解对机器人技术至关重要,能精准刻画运动与场景交互。然而,4D数据集远少于3D数据集,制约了自监督4D模型的扩展性。一种可行方案是将预训练的3D模型迁移至4D感知任务。但实证分析揭示两大瓶颈:过拟合与模态差异。为此,我们提出“对齐后适配”(PointATA)新范式,将参数高效迁移分为两个阶段。第一阶段利用最优传输理论量化3D与4D数据分布差异,训练点对齐嵌入器以缓解模态差距;第二阶段在冻结的3D主干网络中引入高效点视频适配器与空间上下文编码器,增强时序建模能力,抑制过拟合。实验表明,该方法使无时序知识的预训练3D模型以极低参数开销实现动态视频推理,在3D动作识别上达97.21%准确率,4D动作分割提升8.7%,4D语义分割达84.06%,性能媲美甚至超越全微调模型。
原文摘要 · Abstract (English)
Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are far scarcer than 3D ones, which hampers the scalability of self-supervised 4D models. A promising alternative is to transfer 3D pre-trained models to 4D perception tasks. However, rigorous empirical analysis reveals two critical limitations that impede transfer capability: overfitting and the modality gap. To overcome these challenges, we develop a novel "Align then Adapt" (PointATA) paradigm that decomposes parameter-efficient transfer learning into two sequential stages. Optimal-transport theory is employed to quantify the distributional discrepancy between 3D and 4D datasets, enabling our proposed point align embedder to be trained in Stage 1 to alleviate the underlying modality gap. To mitigate overfitting, an efficient point-video adapter and a spatial-context encoder are integrated into the frozen 3D backbone to enhance temporal modeling capacity in Stage 2. Notably, with the above engineering-oriented designs, PointATA enables a pre-trained 3D model without temporal knowledge to reason about dynamic video content at a smaller parameter cost compared to previous work. Extensive experiments show that PointATA can match or even outperform strong full fine-tuning models, whilst enjoying the advantage of parameter efficiency, e.g. 97.21 \% accuracy on 3D action recognition, $+8.7 \%$ on 4 D action segmentation, and 84.06\% on 4D semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。