arXiv:2605.09904cs.CV2026-05

评测视频大模型对物体跨时间保持一致性的能力,发现其仍存在严重短板。

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models

论文配图:TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
图 1 · 摘自论文原文
  • 基于物体轨迹设计多维度时序一致性测试,确保问题依赖连续视觉证据。
  • 筛选出1.79万条需时序推理的问答对,构建2323条人工验证高质量数据集。
  • 揭示当前视频大模型在身份追踪、事件排序和幻觉识别方面存在明显缺陷。

视频大语言模型(Video-LLMs)在通用视频理解方面进展显著,但其保持时间上物体一致性的能力仍缺乏深入研究。现有基准多关注事件识别、动作理解或粗粒度时序推理,很少检验模型在遮挡、消失、重现、状态变化及跨对象交互等情况下能否持续追踪同一物体的身份、状态与连贯性。为此,我们提出TOC-Bench,一个用于评估Video-LLMs时序物体一致性的诊断基准。该基准以物体跟踪为基础,每个被查询主体均关联逐帧轨迹和结构化时序事件时间线。为确保问题必须依赖时序视觉证据而非语言先验、单帧捷径或无序帧线索,设计三层时序必要性过滤协议,剔除60.7%候选问答对,保留17,900条具有时序依赖关系的数据,覆盖10个诊断维度。从中构建包含2,323个高质量问答对的经人工验证基准,覆盖1,951段视频。在代表性Video-LLMs上的实验表明,时序物体一致性仍是未解决的核心挑战,尤其在事件计数、事件排序、身份敏感推理和幻觉感知验证方面表现薄弱,即使在通用视频理解基准上表现良好亦如此。结果表明,以物体为中心的时间连贯性是当前Video-LLMs的关键瓶颈,而TOC-Bench提供了一个聚焦的诊断与改进平台。资源已公开于https://github.com/cjzcjz666/toc_bench.git。

原文摘要 · Abstract (English)

Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored. Existing benchmarks often emphasize event recognition, action understanding, or coarse temporal reasoning, while rarely testing whether models can preserve the identity, state, and continuity of the same object across occlusion, disappearance, reappearance, state transitions, and cross-object interactions. We introduce TOC-Bench, a diagnostic benchmark for evaluating temporal object consistency in Video-LLMs. TOC-Bench is object-track grounded: each queried subject is linked to a per-frame trajectory and a structured temporal event timeline. To ensure that questions require temporally ordered visual evidence rather than language priors, single-frame shortcuts, or unordered frame cues, we design a three-layer temporal-necessity filtering protocol, which removes 60.7% of candidate QA pairs and retains 17,900 temporally dependent items across 10 diagnostic dimensions. From this pool, we construct a human-verified benchmark with 2,323 high-quality QA pairs over 1,951 videos. Experiments on representative Video-LLMs show that temporal object consistency remains a major unsolved challenge, with notable weaknesses in event counting, event ordering, identity-sensitive reasoning, and hallucination-aware verification, even when models perform well on general video understanding benchmarks. These results suggest that object-centric temporal coherence is a key bottleneck for current Video-LLMs, and that TOC-Bench provides a focused platform for diagnosing and improving object-aware temporal reasoning. The resource is available at https://github.com/cjzcjz666/toc_bench.git.

视频理解时序推理物体追踪评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。