构建交通异常推理数据集,推动视觉语言模型从检测到理解的跃迁。
From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

- 基于多尺度视频证据生成结构化事件描述,构建链式思维标注
- 多任务微调使综合得分提升21.4分,超越零样本基线
- 专为交通异常推理设计,适合竞赛与场景理解研究者使用
我们提出TAR(Traffic Anomaly Reasoning)及TAR-Bench数据集,用于训练和评估超越异常检测的视频-语言模型。TAR包含3,670段来自八个公开数据集的监控视频(约26小时),涵盖10项任务,共44,040条链式思维训练标注。其评估部分TAR-Bench包含从17个公开YouTube视频中截取的80个保留片段,配有960条人工标注测试样本。TAR的训练标注通过MAVEN生成,该方法将多尺度视频证据整合为结构化事件描述,再生成问答对与推理轨迹。在TAR-Bench上,11个视觉-语言模型的实验表明,强问答准确率并不能可靠预测时间或场景推理能力。在TAR上进行多任务微调后获得一致性能提升,完整10任务模型的综合得分比零样本基线提高21.4分。TAR与TAR-Bench是AI City Challenge 2026 Track 3的官方训练与域内评估数据集,数据集已发布于https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning。
原文摘要 · Abstract (English)
We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection. TAR contains 44,040 chain-of-thought training annotations across 10 tasks for 3,670 CCTV videos ($\sim$26 hours) from eight public datasets. Its evaluation component, TAR-Bench, contains 960 human-curated test annotations for 80 held-out clips trimmed from 17 public YouTube videos. TAR's training annotations are produced with MAVEN, which consolidates multi-scale video evidence into structured event descriptions before generating question-answer pairs and reasoning traces. On TAR-Bench, eleven vision-language models reveal that strong question-answering accuracy does not reliably predict temporal or scene reasoning ability. Multi-task fine-tuning on TAR yields consistent gains, with the full 10-task model improving aggregate score by 21.4 points over its zero-shot baseline. TAR and TAR-Bench provide the official training and in-domain evaluation data for AI City Challenge 2026 Track 3. The dataset is available at https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。