用专家协作与动态损失提升多模态信息抽取的效率与效果
Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
- 引入通用专家与任务专用专家协同学习跨任务共享与特异性知识
- 在7个数据集上优于传统微调和LoRA方法,参数量相当
- 适合需要高效多任务训练的多模态信息抽取研究者
多模态信息抽取(MIE)从多媒体源中提取结构化信息,近年来受到关注。传统方法分别处理各类任务,错失跨任务知识共享机会。近期方法将任务统一为生成问题,使用带视觉适配器的指令式T5模型,通过全参数微调优化,但计算成本高,且多任务微调常出现梯度冲突,限制性能。为此,我们提出基于成就的多任务损失协同多LoRA专家(C-LoRAE)框架。该方法扩展低秩适配(LoRA)技术,引入通用专家以学习跨MIE任务的共享多模态知识,并设置任务专属专家学习特定指令特征。该设计增强了模型在多个任务间的泛化能力,同时保持任务独立性并缓解梯度冲突。此外,提出基于成就的多任务损失,平衡各任务训练进度,解决因训练样本数量差异导致的不平衡问题。在三个关键MIE任务的七个基准数据集上的实验表明,C-LoRAE在整体性能上优于传统微调和LoRA方法,且训练参数量与LoRA相当。
原文摘要 · Abstract (English)
Multimodal Information Extraction (MIE) has gained attention for extracting structured information from multimedia sources. Traditional methods tackle MIE tasks separately, missing opportunities to share knowledge across tasks. Recent approaches unify these tasks into a generation problem using instruction-based T5 models with visual adaptors, optimized through full-parameter fine-tuning. However, this method is computationally intensive, and multi-task fine-tuning often faces gradient conflicts, limiting performance. To address these challenges, we propose collaborative multi-LoRA experts with achievement-based multi-task loss (C-LoRAE) for MIE tasks. C-LoRAE extends the low-rank adaptation (LoRA) method by incorporating a universal expert to learn shared multimodal knowledge from cross-MIE tasks and task-specific experts to learn specialized instructional task features. This configuration enhances the model's generalization ability across multiple tasks while maintaining the independence of various instruction tasks and mitigating gradient conflicts. Additionally, we propose an achievement-based multi-task loss to balance training progress across tasks, addressing the imbalance caused by varying numbers of training samples in MIE tasks. Experimental results on seven benchmark datasets across three key MIE tasks demonstrate that C-LoRAE achieves superior overall performance compared to traditional fine-tuning methods and LoRA methods while utilizing a comparable number of training parameters to LoRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。