arXiv:2509.08538cs.CVcs.AI2025-09被引 2

MESH benchmark测评大模型视频幻觉,更贴近人类理解方式。

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

  • 用问答形式分层测试物体、特征和动作关系的幻觉
  • 长视频中多主体精细动作识别时幻觉率显著上升
  • 适合评估视频生成与理解模型的真实能力

大型视频模型(LVMs)在大语言模型(LLMs)和视觉模块基础上融合时序信息,以更好理解动态视频内容。尽管取得进展,但容易产生不准确或无关的描述,即幻觉。现有视频幻觉评测依赖人工对视频内容分类,忽视了人类自然理解视频的感知过程。我们提出MESH基准,系统评估LVM中的幻觉问题。MESH采用问答框架,包含二选一和多选题,涵盖目标实例与陷阱实例。它采用自下而上的方法,评估基础物体、粗到细的主体特征以及主体-动作对,符合人类视频理解过程。实验表明,MESH能有效全面识别视频中的幻觉。我们的评估显示,虽然LVMs在识别基础物体和特征方面表现良好,但在处理细粒度信息或较长视频中涉及多个主体的复杂动作时,幻觉概率明显升高。

原文摘要 · Abstract (English)

Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in videos. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos.

视频理解幻觉检测大模型评测问答基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。