arXiv:2503.20633cs.LG2025-03被引 4

用异构专家适配器提升多模态模型微调效率

Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning

  • 设计异构专家适配器,支持多模态专家组合与融合
  • 仅需5-8%参数微调,实现媲美全量微调的性能
  • 适合需要高效多模态微调的研究者和工程师

多模态模型在跨模态任务中表现优异,但因参数量达数十亿而计算成本高昂。参数高效微调(PEFT)通过添加少量可训练组件并冻结预训练参数来缓解此问题。然而,现有方法多聚焦于单模态处理,忽视了多模态任务中的关键模态融合。为此,我们提出异构混合专家适配器,扩展传统PEFT框架以支持多模态专家组合,并增强信息交互。此外,通过改进仿射线性专家设计,在低秩空间实现高效模态融合,仅需5-8%参数微调即达到竞争性性能。在八个下游任务(包括视觉-音频、文本-视觉)上的实验验证了该方法的优越性。

原文摘要 · Abstract (English)

Multi-modal models excel in cross-modal tasks but are computationally expensive due to their billions of parameters. Parameter-efficient fine-tuning (PEFT) offers a solution by adding small trainable components while freezing pre-trained parameters. However, existing methods primarily focus on uni-modal processing, overlooking the critical modal fusion needed for multi-modal tasks. To fill this gap, we propose heterogeneous mixture of experts adapters that extend the traditional PEFT framework to support multi-modal expert combinations and improve information interaction. Additionally, our approach modifies the affine linear expert design to enable efficient modal fusion in a low-rank space, achieving competitive performance with only 5-8\% of the parameters fine-tuned. Experiments across eight downstream tasks, including visual-audio and text-visual, demonstrate the superior performance of the approach.

多模态PEFT专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。