提出新数据集与方法,解决电商多模态模型依赖图片反而失效的问题。
EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models
- 基于40万样本构建多图像电商数据集,评估视觉信息利用效果。
- 实验证明图片未必提升性能,部分场景反而导致模型表现下降。
- 提出SUMEI方法,先预测图像价值再使用,提升模型鲁棒性。
电商平台包含丰富多模态数据,其中多种图像展现产品细节。然而,这些图像是否总能增强理解,还是可能引入冗余甚至降低性能?现有数据集在规模与设计上均有限,难以系统回答此问题。为此,我们提出EcomMMMU,一个包含406,190个样本和8,989,510张图像的电商多模态多任务理解数据集。该数据集由多图像视觉语言数据构成,涵盖8项核心任务,并设有专门的VSS子集,用于评估多模态大模型(MLLMs)有效利用视觉内容的能力。对EcomMMMU的分析显示,产品图像并非始终提升性能,某些情况下反而会降低表现,表明MLLMs在处理丰富视觉内容时存在困难。基于此,我们提出SUMEI——一种数据驱动的方法,通过预判视觉信息效用,有策略地使用多图像进行下游任务。大量实验验证了SUMEI的有效性与鲁棒性。数据与代码已开源:https://github.com/ninglab/EcomMMMU。
原文摘要 · Abstract (English)
E-commerce platforms are rich in multimodal data, featuring a variety of images that depict product details. However, this raises an important question: do these images always enhance product understanding, or can they sometimes introduce redundancy or degrade performance? Existing datasets are limited in both scale and design, making it difficult to systematically examine this question. To this end, we introduce EcomMMMU, an e-commerce multimodal multitask understanding dataset with 406,190 samples and 8,989,510 images. EcomMMMU is comprised of multi-image visual-language data designed with 8 essential tasks and a specialized VSS subset to benchmark the capability of multimodal large language models (MLLMs) to effectively utilize visual content. Analysis on EcomMMMU reveals that product images do not consistently improve performance and can, in some cases, degrade it. This indicates that MLLMs may struggle to effectively leverage rich visual content for e-commerce tasks. Building on these insights, we propose SUMEI, a data-driven method that strategically utilizes multiple images via predicting visual utilities before using them for downstream tasks. Comprehensive experiments demonstrate the effectiveness and robustness of SUMEI. The data and code are available through https://github.com/ninglab/EcomMMMU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。