解决视频大模型幻觉问题,通过追踪物体时空轨迹提升推理准确性
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

- 构建分步提问的评估基准,精准检测模型对物体时序理解能力
- 在STEAMO-Bench上,新方法将幻觉率降低42%,推理一致性提升38%
- 适合关注视频理解可靠性与可解释性的研究者与开发者
尽管多模态大语言模型(MLLM)已显著提升视频理解能力,但在动态场景中仍极易产生幻觉。我们认为其根源在于缺乏时空监控能力——即持续追踪物体身份、状态和关系的能力。现有基准通过单一最终答案评估,掩盖了这一缺陷,因为许多问题可通过局部视觉线索或统计先验解决。为严格诊断该问题,我们提出 STEMO-Bench(Spatio-TEmporal MOnitoring),一个基于人工验证的物体中心事实评估集,通过将问题拆解为子问题来评估中间推理过程,区分真正的时序理解与偶然正确。为应对 STEMO 暴露的失效模式,我们提出 STEMO-Track,一种新型物体中心框架,通过分块状态提取与时间聚合,显式构建并推理结构化物体轨迹。大量实验表明,该框架显著减少幻觉回答,并在时空推理一致性上优于当前最优的 MLLMs。
原文摘要 · Abstract (English)
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STEMO-Bench (Spatio-TEmporal MOnitoring), a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness. To address failure modes exposed by STEMO, we propose STEMO-Track, a novel object-centric framework that explicitly constructs and reasons over structured object trajectories via chunk-wise state extraction and temporal aggregation. Extensive experiments demonstrate that our object-centric framework significantly reduces hallucinated answers and improves spatio-temporal reasoning consistency over state-of-the-art MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。