首个评估长时多模态轨迹推理的基准,挑战大模型跨模态连环推理能力。
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

- 构建覆盖339.5小时长视频的多跳证据链,平均跨度15.1分钟、3.68跳
- 2200道题测试显示最强模型仅达68.29%,开放模型最高51.70%
- 聚焦跨视听流证据整合,适合研究长视频理解与多模态鲁棒性
真实世界音视频理解需串联分散在视听流中的稀疏证据,现有基准多局限于短片段、单模态或单跳感知。我们提出TraceAV-Bench,首个联合评估长时音视频轨迹多跳推理与多模态幻觉鲁棒性的基准。该数据集包含578段长视频(共339.5小时),2200道经严格验证的多选题,覆盖4个评估维度与15个子任务。每道题基于平均3.68跳、跨度15.1分钟的显式多跳证据链。通过三步构建与严苛质检流程生成。对多个代表性OmniLLMs的评测表明,所有模型均面临持续挑战:最强闭源模型(Gemini 3.1 Pro)通用任务准确率仅68.29%,最佳开源模型(Ming-Flash-Omni-2.0)为51.70%,仍有巨大提升空间。分析显示当前模型难以在长时音视频中有效聚合与融合跨模态证据,性能随模态与推理任务变化显著。期望TraceAV-Bench推动OmniLLMs在长时音视频理解上的可靠进展。
原文摘要 · Abstract (English)
Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit multi-hop evidence trajectory that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Further analysis shows that current models struggle to gather and combine evidence across long audio-visual videos, with performance varying across modalities and reasoning tasks. We hope TraceAV-Bench will facilitate progress toward reliable long-form audio-visual reasoning in OmniLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。