arXiv:2602.08355cs.CV2026-02中稿 · ICML被引 2

首个专为电商短视频设计的多模态理解基准,提升商业意图推理能力。

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

  • 构建多模态密度评估框架,量化电商视频复杂度。
  • 推出含1.9万问答对的E-VAds基准,覆盖五类任务。
  • 基于强化学习的模型在少量数据下实现超100%性能提升。

电商短视频是在线视频产业中高收入的细分领域,具有目标导向性强、多模态信号密集的特点。现有模型在此类视频上表现不佳,因主流基准多聚焦通用任务,忽视商业意图推理。本文首先提出多模态信息密度评估框架,量化该领域的复杂性。评估显示,电商内容在视觉、音频和文本模态上的密度显著高于主流数据集,构成更严峻的视频理解挑战。为此,我们推出首个专为电商短视频理解设计的基准——E-VAds,从淘宝精选3,961个高质量视频,涵盖多种商品类别,并利用多智能体系统生成19,785个开放式问答对,包含五类任务。最后,我们开发E-VAds-R1,一种基于强化学习的推理模型,采用多粒度奖励机制MG-GRPO,早期探索阶段提供平滑引导,专家级精度阶段激发非线性激励。实验表明,该模型仅用数百个训练样本,即在商业意图推理上取得109.2%的性能提升。数据已开源:https://github.com/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark。

原文摘要 · Abstract (English)

E-commerce short videos represent a high-revenue segment of the online video industry characterized by a goal-driven format and dense multi-modal signals. Current models often struggle with these videos because existing benchmarks focus primarily on general-purpose tasks and neglect the reasoning of commercial intent. In this work, we first propose a multi-modal information density assessment framework to quantify the complexity of this domain. Our evaluation reveals that e-commerce content exhibits substantially higher density across visual, audio, and textual modalities compared to mainstream datasets, establishing a more challenging frontier for video understanding. To address this gap, we introduce E-commerce Video Ads Benchmark, which is the first benchmark specifically designed for e-commerce short video understanding. We curated 3,961 high-quality videos from Taobao covering a wide range of product categories and used a multi-agent system to generate 19,785 open-ended Q&A pairs, which consist of five distinct tasks. Finally, we develop E-VAds-R1, an RL-based reasoning model featuring a multi-grained reward design called MG-GRPO. This strategy provides smooth guidance for early exploration while creating a non-linear incentive for expert-level precision. Experimental results demonstrate that E-VAds-R1 achieves a 109.2% performance gain in commercial intent reasoning with only a few hundred training samples. Data is available at https://github.com/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark.

电商视频多模态意图推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。