arXiv:2503.18672cs.CV2025-03被引 1

用特征校准与参数融合,让多模态模型持续学习新类别不丢旧知识。

CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion

  • 动态校准视觉特征,平衡原始模型与新任务特征。
  • 基于QR分解融合参数,保持新旧知识平衡,平均准确率超现有方法。
  • 适合需要长期更新的多模态识别系统,如智能监控、跨模态搜索。

随着大规模视觉识别系统中多模态数据的激增,如何在不回看历史数据的前提下持续学习新知识并保留旧知识变得愈发重要。类持续学习(CCL)通过增量引入新类别知识来应对这一挑战,适用于真实世界的大数据应用。传统CCL方法仅依赖视觉特征,而近期基于视觉语言模型(如CLIP)的进展表明,利用预训练多模态知识在CCL中具有巨大潜力。然而,现有方法在缓解灾难性遗忘的同时,难以维持VLM的跨模态泛化能力。为此,我们提出CalFuse框架,通过特征校准与参数融合实现多模态知识的有效整合。CalFuse引入动态特征校准机制,自适应地平衡CLIP原始视觉表示与任务特定特征,既保留模型固有的跨模态泛化能力,又适配新类别。同时,基于QR分解的参数融合策略逐步将新知识与历史任务参数结合,维持学习新类与保留旧知识之间的平衡。大量实验在基准数据集上验证了该方法在大规模多模态持续学习场景下的有效性,其平均准确率和最终任务保留性能均优于现有最先进方法。

原文摘要 · Abstract (English)

With the proliferation of multi-modal data in large-scale visual recognition systems, enabling models to continuously acquire knowledge from evolving data streams while preserving prior information has become increasingly critical. Class-Continual Learning (CCL) addresses this challenge by incrementally incorporating new class knowledge without revisiting historical data, making it essential for real-world big data applications. While traditional CCL methods rely solely on visual features, recent advances in Vision-Language Models (VLMs) such as CLIP demonstrate significant potential for CCL by leveraging pre-trained multi-modal knowledge. However, existing approaches face challenges in mitigating catastrophic forgetting while maintaining the cross-modal generalization capabilities of VLMs. To address these limitations, we propose CalFuse, a framework that synergizes feature Calibration with parameter Fusion to enable effective multi-modal knowledge integration in continual learning scenarios. CalFuse introduces a dynamic feature calibration mechanism that adaptively balances original CLIP visual representations with task-specific features, preserving the model's intrinsic cross-modal generalization while adapting to new classes. Concurrently, a QR decomposition-based parameter fusion strategy progressively integrates newly acquired knowledge with historical task parameters, maintaining equilibrium between learning new class representations and retaining prior knowledge across sequential tasks. Extensive experiments on benchmark datasets validate the effectiveness of our approach in large-scale multi-modal continual learning settings, demonstrating superior performance over state-of-the-art methods in both average accuracy and final task retention.

多模态持续学习视觉语言模型参数融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。