首个评估第一人称视频幻觉的基准,揭示大模型严重误判问题
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- 构建1400段第一人称视频与8000个问答对,专攻视觉听觉幻觉
- 十款大模型平均准确率仅59%,连GPT-4o和Gemini也难幸免
- 开源数据集助力开发更少幻觉的智能视觉系统,适合研究者参考
多模态大语言模型(MLLM)在复杂多模态任务中表现卓越。尽管它们在第三人称和第一人称视频的视觉感知与推理上表现出色,却容易产生幻觉,生成看似合理但不准确的回答。本文提出 EgoIllusion,首个用于评估第一人称视频中 MLLM 幻觉的基准。该基准包含1,400段视频和8,000个由人类标注的开放与封闭式问题,旨在触发第一人称视频中视觉与听觉线索的幻觉。对十种 MLLM 的评估显示显著挑战,包括 GPT-4o 与 Gemini 等强模型,平均准确率仅为59%。EgoIllusion 为开发更鲁棒的 MLLM 基准奠定了基础,并推动更少幻觉的第一人称多模态模型的发展。本基准将开源以保障可复现性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EgoIllusion, a first benchmark to evaluate MLLM hallucinations in egocentric videos. EgoIllusion comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. Evaluations across ten MLLMs reveal significant challenges, including powerful models like GPT-4o and Gemini, achieving only 59% accuracy. EgoIllusion lays the foundation in developing robust benchmarks to evaluate the effectiveness of MLLMs and spurs the development of better egocentric MLLMs with reduced hallucination rates. Our benchmark will be open-sourced for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。