用高质量图文指令数据训练电商多模态模型,提升泛化能力。
Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data
- 构建首个大规模电商多模态指令数据集MMECInstruct
- 提出轻量级框架CASLIE,显著提升图文信息融合效果
- 模型在跨领域场景中表现优异,适合电商应用研究者
利用多模态数据推动电商应用中的突破性进展正受到研究界越来越多关注。然而,基础模型在使用电商多模态数据时面临两大挑战:(1) 缺乏大规模、高质量的多模态基准数据集;(2) 多模态信息融合方法不足。为此,本文首次提出MMECInstruct——首个大规模、高质量的电商多模态指令数据集,并开发了CASLIE框架,一种简单、轻量但高效的电商多模态信息融合方法。基于MMECInstruct,我们在CASLIE框架下微调一系列电商多模态基础模型(简称CASLIE模型)。全面评估表明,这些模型在域内任务上显著优于5类先进基线模型,且在域外设置下仍表现出强泛化能力。MMECInstruct与CASLIE模型已公开,可通过https://ninglab.github.io/CASLIE/获取。
原文摘要 · Abstract (English)
Leveraging multimodal data to drive breakthroughs in e-commerce applications through Multimodal Foundation Models (MFMs) is gaining increasing attention from the research community. However, there are significant challenges that hinder the optimal use of multimodal e-commerce data by foundation models: (1) the scarcity of large-scale, high-quality multimodal benchmark datasets; and (2) the lack of effective multimodal information integration methods. To address these challenges, in this paper, we introduce MMECInstruct, the first-ever, large-scale, and high-quality multimodal instruction dataset for e-commerce. We also develop CASLIE, a simple, lightweight, yet effective framework for integrating multimodal information for e-commerce. Leveraging MMECInstruct, we fine-tune a series of e-commerce MFMs within CASLIE, denoted as CASLIE models. Our comprehensive evaluation demonstrates that CASLIE models substantially outperform 5 categories of advanced baseline models in the in-domain evaluation. Moreover, CASLIE models show strong generalizability to out-of-domain settings. MMECInstruct and CASLIE models are publicly accessible through https://ninglab.github.io/CASLIE/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。