arXiv:2508.11801cs.CVcs.CL2025-08中稿 · CIKM 2025被引 1

首个视频转文本的电商属性值抽取数据集,支持14个品类172项属性。

VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models

  • 构建跨14类商品的视频-文本属性值匹配数据集
  • 通过CLIP-MoE过滤后保留22.4万训练/2.5万测试样本
  • 揭示视频模型在开放场景下仍存在显著性能瓶颈

属性值抽取(AVE)对电商平台结构化商品信息至关重要。然而现有数据集主要局限于文本或图像到文本的场景,缺乏视频支持、属性覆盖不全且未公开。为此,我们提出VideoAVE,首个面向电商领域的公开视频-文本属性值抽取数据集,涵盖14个不同品类和172个独特属性。为保障数据质量,我们设计基于CLIP的后处理专家混合过滤系统(CLIP-MoE),剔除视频与商品不匹配的样本,最终获得22.4万条训练数据和2.5万条评估数据。为进一步评估数据集可用性,我们建立全面基准,评估多个前沿视频视觉语言模型(VLMs)在属性条件值预测与开放属性-值对提取任务中的表现。结果表明,视频到文本的AVE仍是极具挑战的问题,尤其在开放设置中,现有模型尚无法有效利用时序信息,亟需更先进的VLMs。VideoAVE数据集与基准代码已开源:https://github.com/gjiaying/VideoAVE。

原文摘要 · Abstract (English)

Attribute Value Extraction (AVE) is important for structuring product information in e-commerce. However, existing AVE datasets are primarily limited to text-to-text or image-to-text settings, lacking support for product videos, diverse attribute coverage, and public availability. To address these gaps, we introduce VideoAVE, the first publicly available video-to-text e-commerce AVE dataset across 14 different domains and covering 172 unique attributes. To ensure data quality, we propose a post-hoc CLIP-based Mixture of Experts filtering system (CLIP-MoE) to remove the mismatched video-product pairs, resulting in a refined dataset of 224k training data and 25k evaluation data. In order to evaluate the usability of the dataset, we further establish a comprehensive benchmark by evaluating several state-of-the-art video vision language models (VLMs) under both attribute-conditioned value prediction and open attribute-value pair extraction tasks. Our results analysis reveals that video-to-text AVE remains a challenging problem, particularly in open settings, and there is still room for developing more advanced VLMs capable of leveraging effective temporal information. The dataset and benchmark code for VideoAVE are available at: https://github.com/gjiaying/VideoAVE

视频理解属性抽取多模态电商

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。