arXiv:2506.02448cs.CVcs.AI2025-06AAAI被引 10

构建大规模视频事件理解数据集,助力AI解析动态事件演变。

VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos

  • 基于电影解说视频构建事件结构与逻辑关系标注体系
  • 涵盖超2.3万条事件,覆盖广泛语义层级与动态演化模式
  • 提供基准模型与公开数据,适合视频理解与认知建模研究者

尽管视觉事件对人类认知有重要影响,但其复杂的结构、语义层次和动态演变使得视频事件理解对AI而言仍是挑战。为此,我们提出视频事件理解任务,旨在从视频中提取事件脚本并基于脚本进行预测。为此,我们引入了VidEvent——一个包含超过23,000条高质量标注事件的大规模数据集,事件结构细致、语义层级丰富,并从电影解说视频中提取了逻辑关系。数据集通过严谨的标注流程构建,确保高质量与可靠性。我们还提供了全面的基线模型,详述其架构与性能指标,作为未来研究的基准,支持对比与改进。对VidEvent及基线模型的分析表明,该数据集具有推动视频事件理解发展的潜力,并鼓励探索创新算法。数据集及相关资源已公开:www.videvent.top。

原文摘要 · Abstract (English)

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose the task of video event understanding that extracts event scripts and makes predictions with these scripts from videos. To support this task, we introduce VidEvent, a large-scale dataset containing over 23,000 well-labeled events, featuring detailed event structures, broad hierarchies, and logical relations extracted from movie recap videos. The dataset was created through a meticulous annotation process, ensuring high-quality and reliable event data. We also provide comprehensive baseline models offering detailed descriptions of their architecture and performance metrics. These models serve as benchmarks for future research, facilitating comparisons and improvements. Our analysis of VidEvent and the baseline models highlights the dataset's potential to advance video event understanding and encourages the exploration of innovative algorithms and models. The dataset and related resources are publicly available at www.videvent.top.

视频理解事件检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。