构建大规模多任务多主题图文实体链接数据集,提升模型泛化能力
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
- 提出构建流程,生成覆盖9类任务、5类主题的7.9万条多模态数据
- 基于新数据微调模型后,准确率提升9.3%至25%,显著优于基线
- 适合研究多模态理解、实体对齐与跨模态训练的开发者使用
多模态实体链接(MEL)是众多下游任务的基础。现有MEL数据集存在规模小、主题类型少、任务覆盖有限等问题,难以有效提升多模态模型的链接能力。为此,本文提出数据构建流程,发布大规模数据集$M^3EL$,包含79,625个实例,覆盖9种多样化的多模态任务和5种不同主题。为增强模型对多模态任务的适应性,提出模态增强训练策略。以$M^3EL$为语料,基于$ extit{CLIP}( extit{ViT}- extit{B}- extit{32})$训练$ extit{CLIP}_{ extit{ND}}$模型,并与现有基线对比。实验显示,现有模型表现远低于预期(准确率49.4%-75.8%)。分析表明,小规模数据、模态任务覆盖不足及主题多样性差导致模型泛化能力弱。$M^3EL$有效缓解上述问题,经$M^3EL$微调的$ extit{CLIP}_{ extit{ND}}$在多个任务上平均提升9.3%至25%。数据集已公开于https://anonymous.4open.science/r/M3EL。
原文摘要 · Abstract (English)
Multi-modal Entity Linking (MEL) is a fundamental component for various downstream tasks. However, existing MEL datasets suffer from small scale, scarcity of topic types and limited coverage of tasks, making them incapable of effectively enhancing the entity linking capabilities of multi-modal models. To address these obstacles, we propose a dataset construction pipeline and publish $M^3EL$, a large-scale dataset for MEL. $M^3EL$ includes 79,625 instances, covering 9 diverse multi-modal tasks, and 5 different topics. In addition, to further improve the model's adaptability to multi-modal tasks, We propose a modality-augmented training strategy. Utilizing $M^3EL$ as a corpus, train the $\textit{CLIP}_{\textit{ND}}$ model based on $\textit{CLIP} (\textit{ViT}-\textit{B}-\textit{32})$, and conduct a comparative analysis with an existing multi-modal baselines. Experimental results show that the existing models perform far below expectations (ACC of 49.4%-75.8%), After analysis, it was obtained that small dataset sizes, insufficient modality task coverage, and limited topic diversity resulted in poor generalisation of multi-modal models. Our dataset effectively addresses these issues, and the $\textit{CLIP}_{\textit{ND}}$ model fine-tuned with $M^3EL$ shows a significant improvement in accuracy, with an average improvement of 9.3% to 25% across various tasks. Our dataset is available at https://anonymous.4open.science/r/M3EL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。