arXiv:2607.01117cs.CV2026-07

评测视频大模型对运动幻觉的敏感度,揭示其推理偏差根源。

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

论文配图:MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
图 1 · 摘自论文原文
  • 构建多源幻觉诊断基准,涵盖共现先验、序列推断等三类问题。
  • 10个模型测试显示,动作识别强的模型反而更易产生运动幻觉。
  • 发现序列推理导致幻觉最严重,模型常凭部分动作推测完整动作。

视频大语言模型(VideoLLMs)在视频理解方面已取得显著进展,但仍存在与视觉证据不符的幻觉问题。现有基准主要关注物体幻觉或粗粒度动作感知,忽视了关键的视频特有问题:运动幻觉,即模型推断出视频中不存在的人体动作。本文提出MoHallBench,一个用于诊断视频大模型运动幻觉的基准。该基准系统评估三大幻觉来源:共现先验、序列推理和相似性混淆。包含11,306个视频片段和40,493个问答对,覆盖二选一、多选和生成式三种设置。我们还引入双向提问协议与偏差感知指标,以减少二选一评估中的肯定偏见。对十款近期开源的VideoLLMs实验表明,动作识别能力与抗幻觉能力明显解耦:在正样本上表现优异的模型,在对抗性负样本上却频繁失败。其中,序列推理引起的幻觉最为严重,说明当前模型容易从部分运动线索过度推断预期结果。分析进一步证实,更强的先验知识和更细粒度的相似性匹配会显著加剧幻觉。我们希望MoHallBench能推动未来对视频大模型运动幻觉的评估与缓解。

原文摘要 · Abstract (English)

Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object hallucination or coarse action perception, leaving a key video-specific problem underexplored: motion hallucination, in which models infer human motions that are absent from the video. We present MoHallBench, a benchmark for diagnosing motion hallucination in VideoLLMs. MoHallBench systematically evaluates three major sources of hallucination: co-occurrence priors, sequential inference, and similarity confusion. It contains 11,306 video clips and 40,493 question-answer pairs, covering binary-choice, multiple-choice, and generative settings. We further introduce a bi-directional questioning protocol with bias-aware metrics to reduce affirmation bias in binary evaluation. Experiments on ten recent open-source VideoLLMs reveal a clear decoupling between action recognition and hallucination resistance, as models that perform well on positive action recognition often fail on adversarial negatives. Among all settings, sequential inference hallucination is the most severe, showing that current models tend to over-infer expected outcomes from partial motion cues. Our analyses further confirm that stronger priors and finer-grained similarity substantially amplify hallucination. We hope MoHallBench can facilitate future evaluation and mitigation of motion hallucination in VideoLLMs.

视频理解幻觉检测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。