arXiv:2608.01948cs.CV2026-08

构建大规模仿真事件数据集,推动长时序动作理解研究。

Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding

论文配图:Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding
图 1 · 摘自论文原文
  • 从原始视频生成事件体素,保留时间映射与重建质量
  • 在长视频上实现66.42%动作识别准确率,定位mAP达29.0
  • 适合研究事件流建模、多模态对齐及在线动作定位的团队

长时序事件驱动的动作理解因现有数据集多为短裁剪片段而发展受限,且真实事件流密集时间标注成本高昂。本文提出Event ActivityNet,一个基于人工标注未剪辑ActivityNet视频生成的大规模仿真事件基准。包含3,263个视频、200类动作、总时长106.94小时,配有5-bin和9-bin事件体素表示、动作时间标注与时间戳字幕。该基准支持标注段落动作识别、辅助事件-语言对齐及因果在线时间动作定位。事件体素直接由解码帧序的非插值源视频生成,保留每视频的名义或平均帧率元数据以近似时间映射,并使用动作中心重建LPIPS作为可重建内容的软诊断。建立自适应事件分帧、提示-字幕对齐、纯事件、纯RGB及混合模式定位基线。在渐进式嵌套尺度训练下,识别Top-1准确率从52.25提升至66.42,线上定位平均mAP从21.7升至29.0。阶段性预训练+原生事件微调在多种监督预算下均优于仅目标训练与联合从头训练。该基准为长时序事件建模提供可扩展平台,但原相机评估仍对部署结论至关重要。

原文摘要 · Abstract (English)

Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting native event streams with dense temporal annotations is costly. We introduce Event ActivityNet, a large-scale simulated-event benchmark derived from human-annotated, untrimmed ActivityNet videos. It comprises 3,263 videos, 200 action classes, and 106.94 hours, with matched 5-bin and 9-bin event-voxel representations, temporal action annotations, and timestamped captions. The benchmark supports annotated-segment action recognition, auxiliary event-language alignment, and causal online temporal action localization. We generate event voxels directly from non-interpolated source videos in decoded frame order, retain per-video rational nominal or average frame-rate metadata for approximate time mapping, and use action-center reconstruction LPIPS as a soft diagnostic of retained reconstructable content. We establish baselines for adaptive event framing, prompt-caption alignment, and event-only, RGB-only, and RGB-event localization. Under a progressive nested-scale training protocol, recognition Top-1 accuracy increases from 52.25 to 66.42, while online temporal localization average mAP improves from 21.7 to 29.0. Moreover, staged Event ActivityNet pretraining followed by native-event fine-tuning consistently outperforms target-only and joint-from-scratch training across multiple supervision budgets. Event ActivityNet provides a scalable benchmark for long-horizon event modeling, although native-camera evaluation remains essential for deployment-oriented conclusions.

事件视觉动作识别长时序建模多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。