系统梳理视频大模型幻觉问题,分类并提出解决方向。
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
- 将幻觉分为动态扭曲和内容虚构两类,每类含两个子类型。
- 指出时间建模能力弱与视觉定位不足是主要成因。
- 适合研究视频多模态模型可靠性与评估的学者参考。
尽管视频-语言建模取得显著进展,视频大语言模型(Vid-LLMs)中的幻觉问题仍持续存在,表现为输出看似合理却与输入视频内容矛盾。本综述对Vid-LLMs中的幻觉进行了系统分析,提出一个分类体系,将其分为两大核心类型:动态扭曲与内容虚构,每类包含两个子类型,并列举代表性案例。基于此分类,回顾了近期在幻觉评估与缓解方面的进展,涵盖关键基准、度量指标与干预策略。进一步分析了动态扭曲与内容虚构的根本原因,通常源于时间表征能力有限及视觉定位不足。这些洞察为未来工作指明方向,包括开发具备运动感知能力的视觉编码器与整合反事实学习技术。本综述整合了零散的研究成果,旨在推动对Vid-LLMs中幻觉问题的系统理解,为构建鲁棒可靠的视频-语言系统奠定基础。相关文献更新列表见https://github.com/hukcc/Awesome-Video-Hallucination。
原文摘要 · Abstract (English)
Despite significant progress in video-language modeling, hallucinations remain a persistent challenge in Video Large Language Models (Vid-LLMs), referring to outputs that appear plausible yet contradict the content of the input video. This survey presents a comprehensive analysis of hallucinations in Vid-LLMs and introduces a systematic taxonomy that categorizes them into two core types: dynamic distortion and content fabrication, each comprising two subtypes with representative cases. Building on this taxonomy, we review recent advances in the evaluation and mitigation of hallucinations, covering key benchmarks, metrics, and intervention strategies. We further analyze the root causes of dynamic distortion and content fabrication, which often result from limited capacity for temporal representation and insufficient visual grounding. These insights inform several promising directions for future work, including the development of motion-aware visual encoders and the integration of counterfactual learning techniques. This survey consolidates scattered progress to foster a systematic understanding of hallucinations in Vid-LLMs, laying the groundwork for building robust and reliable video-language systems. An up-to-date curated list of related works is maintained at https://github.com/hukcc/Awesome-Video-Hallucination .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。