arXiv:2606.23061cs.CVcs.AI2026-06

提出新基准诊断视频运动推理中的幻觉问题,提升纠正指令准确性。

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

论文配图:MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
图 1 · 摘自论文原文
  • 构建1540个细粒度问题的基准,评估方向、归属、时间三类运动幻觉。
  • 现有大模型在跨视频对比中幻觉率高,平均性能不足基准水平。
  • 引入测量注入法,无需训练即可提升10.6%准确率,适合视频理解研究者。

跨视频对比中的运动指令生成旨在生成描述查询与参考动作差异的修正反馈。然而,现有模型常产生不反映真实运动学差异的幻觉指令。为此,我们提出MotionHalluc,一个专门用于评估配对视频对比中运动幻觉的基准,包含553对视频上的1540个细粒度问题,从方向性、归属性和时间性三个核心维度评估幻觉。对前沿多模态大模型的广泛评估显示其高度易受此类幻觉影响。我们进一步提出感知-解析-验证(PPV)方法,作为无训练的测量提取与验证基线,将候选指令转化为可执行的测量查询,并在推理时提供运动学测量值。结果表明,该简单测量注入方法在多个模型上平均提升10.6%性能,表明显式量化测量对减少跨视频对比中的幻觉至关重要。代码与数据集将在论文接受后公开。

原文摘要 · Abstract (English)

Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing models often generate instructions that exhibit motion hallucinations, failing to reflect actual kinematic differences between paired videos. To systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion hallucinations in paired-video comparison. MotionHalluc comprises 1540 fine-grained questions over 553 video pairs, evaluating hallucinations along three core dimensions: (1)directional hallucination, (2)attributional hallucination, and (3)temporal hallucination. Extensive evaluations of state-of-the-art large multimodal models demonstrate high susceptibility to these hallucinations. Furthermore, we provide Perceive-Parse-Verify (PPV) as a training-free measurements extraction and verification baseline that converts candidate instructions into executable measurement queries and supplies kinematic measurements at inference time. Our results show that this simple measurements injection yields an average 10.6% performance gain across models, suggesting that motion reasoning with explicit quantitative measurements is a key factor in reducing hallucinations in cross-video comparison. Our code and dataset will be made publicly available upon acceptance.

视频理解运动推理幻觉检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。