arXiv:2504.08222cs.CVcs.AI2025-04ICLR被引 8

构建首个聚焦快速频繁细粒度事件的视频分析基准数据集

F$^3$Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos

  • 设计涵盖千类事件、带精确时间戳的多粒度视频数据集
  • 现有方法在该数据集上表现显著不足,验证任务难度
  • 提出新模型F$^3$ED,提升细粒度事件检测性能

快速、频繁、细粒度(F$^3$)事件分析在视频理解与多模态大模型中面临巨大挑战。现有方法因运动模糊和细微视觉差异难以高精度识别满足所有F$^3$特征的事件。为此,我们推出F$^3$Set,一个用于精准F$^3$事件检测的基准数据集。该数据集具有大规模与高细节特征,通常包含超过1,000种事件类型,配有精确时间戳,并支持多层级粒度。目前F$^3$Set涵盖多个体育赛事数据集,未来可扩展至其他应用。我们在F$^3$Set上评估了主流时序动作理解方法,揭示现有技术存在显著瓶颈。此外,我们提出F$^3$ED新方法,在F$^3$事件检测中取得更优效果。数据集、模型及评测代码已公开于https://github.com/F3Set/F3Set。

原文摘要 · Abstract (English)

Analyzing Fast, Frequent, and Fine-grained (F$^3$) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F$^3$ criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F$^3$Set, a benchmark that consists of video datasets for precise F$^3$ event detection. Datasets in F$^3$Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently, F$^3$Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F$^3$Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F$^3$ED, for F$^3$ event detections, achieving superior performance. The dataset, model, and benchmark code are available at https://github.com/F3Set/F3Set.

视频分析事件检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。