构建首个聚焦异常实例跟踪的视频异常理解评测基准
TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

- 以事件轨迹为中心设计联合评估框架
- 包含1118段视频、1454条轨迹和20万+像素级掩码
- 适合研究视频异常理解与视觉定位的学者
人类通过连贯的感知过程理解异常事件:识别焦点实例,追踪其行为发展,并解释为何违反场景预期。视频异常理解(VAU)旨在赋予模型类似能力,从判断视频是否异常转向解释事件如何演变及为何重要。尽管近期视觉-语言模型(VLMs)能生成详细且合理的异常描述,但其语义流畅性未必保证解释始终与正确实例在时间上对齐。现有基准通常分协议评估跟踪与语义理解,未能衡量这种实例-语义不一致。为此,我们提出TAU-Bench,一个以轨迹为中心的联合评测基准,涵盖1,118个视频、1,454条轨迹、202,438个像素级掩码,覆盖49类事件和45类场景,并提供连接实例识别、事件理解与场景推理的轨迹级标注。为规模化构建,我们开发了集成异常适配性过滤、实例轨迹构建、层次化字幕标注与人工质量控制的数据引擎。对代表性VLM家族的评估显示,即使生成合理解释的模型,也可能无法可靠定位和追踪正确实例,揭示了语义推理与视觉定位间的持续差距。该发现强调了基于实例的评估对实现更真实可靠的VAU系统的重要性。
原文摘要 · Abstract (English)
Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision--language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance--semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。