arXiv:2601.15643cs.CV2026-01

提出统一多模态持续感知框架,解决增量学习中的遗忘与语义混淆问题。

Evolving Without Ending: Unifying Multimodal Incremental Learning for Continual Panoptic Perception

  • 设计跨模态协同编码器与可塑知识继承模块,实现多模态特征融合与记忆保留。
  • 在多个多模态数据集上达到领先性能,尤其在细粒度增量任务中提升显著。
  • 无需示例回放即可持续进化,适合长期运行的智能感知系统部署。

持续学习(CL)是构建智能感知AI系统的重要方向,但现有研究主要聚焦单任务场景,难以应对多任务与多模态挑战。除灾难性遗忘外,多任务持续学习还导致跨模态对齐中的语义混淆,造成模型性能严重下降。本文将持续学习扩展至持续全景感知(CPP),融合多模态与多任务持续学习,通过像素级、实例级与图像级联合解析提升整体感知能力。我们形式化了多模态场景下的持续学习任务,并提出端到端的持续全景感知模型。该模型包含协同跨模态编码器(CCE)用于多模态嵌入表示;通过对比特征蒸馏与实例蒸馏设计可塑知识继承模块,缓解任务交互引发的遗忘问题。同时引入跨模态一致性约束,构建CPP+框架,确保多任务增量更新时的语义对齐。此外,采用非对称伪标签机制,实现无需示例回放的模型持续演进。在多个多模态数据集及多样化持续学习任务上的大量实验表明,所提模型在细粒度任务中表现尤为突出。

原文摘要 · Abstract (English)

Continual learning (CL) is a great endeavour in developing intelligent perception AI systems. However, the pioneer research has predominantly focus on single-task CL, which restricts the potential in multi-task and multimodal scenarios. Beyond the well-known issue of catastrophic forgetting, the multi-task CL also brings semantic obfuscation across multimodal alignment, leading to severe model degradation during incremental training steps. In this paper, we extend CL to continual panoptic perception (CPP), integrating multimodal and multi-task CL to enhance comprehensive image perception through pixel-level, instance-level, and image-level joint interpretation. We formalize the CL task in multimodal scenarios and propose an end-to-end continual panoptic perception model. Concretely, CPP model features a collaborative cross-modal encoder (CCE) for multimodal embedding. We also propose a malleable knowledge inheritance module via contrastive feature distillation and instance distillation, addressing catastrophic forgetting from task-interactive boosting manner. Furthermore, we propose a cross-modal consistency constraint and develop CPP+, ensuring multimodal semantic alignment for model updating under multi-task incremental scenarios. Additionally, our proposed model incorporates an asymmetric pseudo-labeling manner, enabling model evolving without exemplar replay. Extensive experiments on multimodal datasets and diverse CL tasks demonstrate the superiority of the proposed model, particularly in fine-grained CL tasks.

持续学习多模态全景分割知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。