arXiv:2609.07009cs.LGcs.CV2026-09

解决多模态持续学习中模态组合动态变化的难题

NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts

论文配图:NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts
图 1 · 摘自论文原文
  • 用专家组合机制灵活应对不同模态任务
  • 在4个真实数据集上显著优于现有方法
  • 适合研究动态多模态智能系统的学者

多模态持续学习有望推动具备类人智能的智能体发展,通过跨模态持续学习新任务。然而,现有方法通常假设每项任务的模态集合是预定义且固定的。本文探讨更现实的学习场景——动态多模态持续学习,其中任务间的模态集合可变。该设置面临两大挑战:(i) 空间-时间灾难性遗忘,(ii) 自适应多模态融合。为此,我们提出NeuCME(神经专家组合框架),包含三个核心组件:模态组合回放、多门控专家混合模型与任务相关性引导的蒸馏机制。此外,我们设计评估指标量化任务序列的动态性,并建立涵盖不同动态程度的综合性基准。在四个真实世界数据集上的大量实验表明,NeuCME显著优于当前最优方法。

原文摘要 · Abstract (English)

Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more realistic learning setting, referred to as dynamic multimodal continual learning, in which the set of modalities may vary across tasks rather than remaining fixed. This setting involves two primary challenges: (i) spatio-temporal catastrophic forgetting and (ii) adaptive multimodal fusion. To address these challenges, we propose NeuCME (as shorthand for \textbf{Neu}ral \textbf{C}ombinatorics of \textbf{M}ultiple \textbf{E}xperts), a novel framework designed to effectively learn and integrate knowledge across tasks with varying modalities. The proposed NeuCME model comprises three key components, namely modality-combinational rehearsal, multi-gated mixture-of-experts, and task relevance-guided distillation. Furthermore, we formulate an evaluation metric to quantify the dynamism of task sequences and then set up a comprehensive benchmark with different degrees of dynamism. Extensive experiments using four real-world datasets demonstrate that the proposed NeuCME outperforms state-of-the-art methods markedly.

多模态学习持续学习专家混合动态任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。