arXiv:2504.08422cs.CV2025-04

解决视觉与点云跨模态增量学习的遗忘问题,提升机器人动态感知能力。

CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning

  • 用对比学习预训练,融合掩码点云与多视角图像增强跨模态对齐
  • 冻结主干网络,通过原型对齐实现旧类别知识保留与新类别持续学习
  • 首个跨模态图像-点云增量学习基准,适合机器人持续学习研究者

图像-点云增量学习有助于3D视觉机器人从2D图像中持续学习类别知识,提升在动态环境中的感知能力。然而,现有方法或仅缓解单模态遗忘,或假设训练与测试数据间无模态差异,无法应对跨模态场景。本文首次探索该任务,提出基准CMIP-CIL,并缓解跨模态灾难性遗忘。预训练阶段采用掩码点云与渲染多视角图像,在对比学习框架下增强视觉模型对图像-点云对应关系的泛化能力。增量阶段冻结主干网络,推动物体表示靠近各自原型,有效保留并泛化旧类别知识,同时学习新类别。在基准数据集上开展全面实验,结果表明该方法达到领先性能,显著优于基线方法。

原文摘要 · Abstract (English)

Image-point class incremental learning helps the 3D-points-vision robots continually learn category knowledge from 2D images, improving their perceptual capability in dynamic environments. However, some incremental learning methods address unimodal forgetting but fail in cross-modal cases, while others handle modal differences within training/testing datasets but assume no modal gaps between them. We first explore this cross-modal task, proposing a benchmark CMIP-CIL and relieving the cross-modal catastrophic forgetting problem. It employs masked point clouds and rendered multi-view images within a contrastive learning framework in pre-training, empowering the vision model with the generalizations of image-point correspondence. In the incremental stage, by freezing the backbone and promoting object representations close to their respective prototypes, the model effectively retains and generalizes knowledge across previously seen categories while continuing to learn new ones. We conduct comprehensive experiments on the benchmark datasets. Experiments prove that our method achieves state-of-the-art results, outperforming the baseline methods by a large margin.

增量学习跨模态点云视觉机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。