arXiv:2607.12112cs.LGcs.AI2026-07

解决联邦多模态模型持续学习中的遗忘问题,提升系统长期可靠性。

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

论文配图:Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
图 1 · 摘自论文原文
  • 分模态弹性正则化保护视觉、语言和跨模态表征
  • 客户端生成嵌入级重放数据,无需共享原始数据
  • 基于任务相似性的梯度聚合稳定全局学习轨迹

在分布式网络中对多模态大语言模型(MLLM)进行联邦微调,可在保护隐私的同时适应动态数据流,但核心障碍是灾难性遗忘——连续任务更新会抹除先前习得的视觉、语言及跨模态知识。这一问题在内容审核等安全敏感场景中尤为关键。为此,我们提出联邦持续多模态学习(FedCMM)框架,将持续学习机制融入联邦优化流程的三个层次:参数层面,采用分模态弹性权重融合,为视觉编码器、语言主干和跨模态投影器分别计算费舍尔信息矩阵,实现模态特异性保护;数据层面,各客户端训练轻量级本地生成重放模块,无须共享原始数据即可合成嵌入级多模态重放样本;聚合层面,基于任务相似性的梯度聚合自动过滤并重加权客户端更新,通过梯度余弦相似性抑制冲突方向,稳定全局学习路径。在两个基准上的实验表明,FedCMM在准确率和反向迁移性能上均优于近期基线,验证了全链路、分模态优化在异构网络部署中实现鲁棒演化适应的有效性。

原文摘要 · Abstract (English)

Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations. Addressing this challenge is especially critical for autonomous networked AI operating in safety-sensitive domains, such as content moderation, where reliable retention of prior knowledge underpins system integrity. To overcome this, we propose Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels. At the parameter level, modality-aware elastic weight consolidation computes separate Fisher information matrices for the vision encoder, language backbone, and cross-modal projector, providing granular, asymmetry-aware protection against modality-specific forgetting. At the data level, each client trains a lightweight local generative replay module to synthesize raw-data-free embedding-level multimodal replay tuples without any raw data sharing. At the aggregation level, Task-similarity-aware gradient aggregation autonomously filters and reweights client updates by gradient cosine similarity, suppressing conflicting directions and stabilizing the global learning trajectory. Extensive experiments on two benchmarks demonstrate that FedCMM consistently outperforms recent baselines on accuracy and backward transfer, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.

联邦学习持续学习多模态遗忘抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。