arXiv:2411.14717cs.LGcs.CL2024-11被引 13

解决多模态数据异构下的联邦微调难题,提升模型实用性

FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

  • 设计通用框架融合经典联邦学习与无模态依赖策略
  • 在五大数据集上实现跨域多模态异构场景下性能提升
  • 适合关注隐私保护与多源异构数据融合的研究者

多模态大语言模型(MLLMs)在处理和理解多模态数据方面取得了显著进展。通过联邦学习(FL)对MLLMs进行微调,可引入私有数据源以扩大训练数据范围,从而增强其在隐私敏感领域的实际应用能力。然而,现有研究仍处于初级阶段,尤其缺乏对真实应用场景中多模态异构性的系统应对。本文提出一个基准测试,用于评估不同多模态异构场景下MLLM联邦微调的性能,为该领域未来研究奠定基础。该基准包含两个轻量级MLLM、两个下游任务、三个评估指标及五个跨三领域数据集,并涵盖六种对比基线,覆盖超过十类模态异构类型,涉及四种多模态场景。为应对多模态异构挑战,我们构建了通用的FedMLLM框架,整合经典联邦学习方法与两种无模态依赖策略。大量实验表明,所提出的联邦范式通过扩展训练数据范围并缓解多模态异构问题,显著提升了MLLM性能。代码见补充材料。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs with Federated Learning (FL) allows for expanding the training data scope by including private data sources, thereby enhancing their practical applicability in privacy-sensitive domains. However, current research remains in the early stage, particularly in addressing the \textbf{multimodal heterogeneities} in real-world applications. In this paper, we introduce a benchmark to evaluate the performance of federated fine-tuning of MLLMs across various multimodal heterogeneous scenarios, laying the groundwork for future research in the field. Our benchmark includes two lightweight MLLMs, two downstream tasks, three evaluation metrics, and five datasets across three domains, along with six comparison baselines, covering over ten types of modality heterogeneities across four multimodal scenarios. To address the challenges posed by multimodal heterogeneity, we develop a general FedMLLM framework that integrates classic FL methods alongside two modality-agnostic strategies. Extensive experimental results show that our proposed FL paradigm improves the performance of MLLMs by broadening the range of training data and mitigating multimodal heterogeneity. Code is available in supplementary materials.

联邦学习多模态大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。