首个面向长视频多事件检索的大规模数据集,提升真实场景下的图文匹配能力。
MELON: A Large-Scale Dataset for Multi-Event Text-to-Long-Video Retrieval

- 构建包含多事件标注的长视频数据集,支持复杂语义理解
- 提出多事件感知损失,使模型能区分完整与部分事件匹配
- 适合研究长视频理解、多事件检索的学者和工程师
现有文本-视频检索数据集主要由包含单一主导事件的短片段构成,虽适用于基础视觉语言对齐评估,但难以反映真实检索场景。在真实场景中,长视频自然包含多个语义上不同的事件,而单个文本查询可能对应多个非连续的时间段。为弥合这一差距,我们提出MELON,首个专为长视频多事件结构设计的大规模数据集。MELON为每段视频显式标注多个事件区间及其对应文本描述,支持长视频中多事件理解的训练与评估。此外,我们提出一种多事件感知损失,促使模型区分完整事件与部分事件匹配,显著提升检索准确率。MELON数据集与所提损失共同构建了扩展文本到视频检索至复杂长视频场景的坚实基础,为该领域未来研究提供更真实的评估环境。
原文摘要 · Abstract (English)
Existing text-video retrieval datasets primarily consist of short-form clips containing a single dominant event. While suitable for measuring basic vision-language alignment, they are limited in capturing real-world retrieval scenarios, where long-form videos naturally contain multiple semantically distinct events and a single text query may correspond to several non-contiguous temporal segments. To bridge this gap, we introduce MELON, the first large-scale dataset designed to extend text-video retrieval to long-form videos featuring complex, multi-event structures. MELON explicitly annotates multiple event intervals per video along with their corresponding textual descriptions, enabling both training and evaluation of multi-event understanding in long, untrimmed videos. In addition, we propose a multi-event aware loss that encourages models to differentiate between full-event and partial-event matches, yielding substantial improvements in retrieval accuracy. Together, the MELON dataset and our proposed loss establish a robust foundation for expanding text-to-video retrieval to complex long-form scenarios and provide a more realistic evaluation setting for future research in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。