新基准测试视频模型在背景一致时的幻觉问题,更精准评估理解能力。
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs

- 用背景一致但前景不同的视频对,隔离出真实幻觉误差
- 包含1000组对抗性视频对和1.1万条时空问答数据
- 适合研究视频模型鲁棒性与幻觉检测的学者使用
我们提出VidPair-Halluc,一个用于评估大视频模型(LVMs)视频幻觉的新基准,具备严格受控条件。与以往依赖文本扰动或对抗性提问、忽略视觉背景一致性的方法不同,VidPair-Halluc采用背景高度相似但前景语义迥异的视频对,使模型错误可准确归因于真实幻觉而非背景差异。该基准通过PairFlow流程构建,结合近期文本到图像与视频生成技术,系统化生成故事、生成连贯视频片段,并组合成对抗性视频对。涵盖十个语义维度的空间与时间推理,共包含1000组高质量对抗性视频对及11000条时空问答对,实现对背景与前景变化的精确控制。主流LVMs的评估显示其在对抗环境下仍难以实现稳健的细粒度视频理解。代码与数据已公开于https://jethrojames.github.io/VidPair-Halluc/。
原文摘要 · Abstract (English)
We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text-based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair-Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text-to-image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair-Halluc comprises 1K high-quality adversarial video pairs and 11K spatio-temporal QA pairs with control over background and foreground variations. Evaluations on mainstream LVMs show persistent difficulty with robust fine-grained video understanding in adversarial settings, and code and data are available at the https://jethrojames.github.io/VidPair-Halluc/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。