首个音视频大模型幻觉评测基准,揭示多模态模型易出错的根源。
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

- 构建跨模态幻觉评测框架,测试音视频理解与推理能力
- 多数音视频大模型在跨模态交互中幻觉率超40%
- 训练数据可显著提升模型抗幻觉能力,适合多模态研究者
随着大语言模型的成功,将其拓展至新模态成为多模态理解的重要范式。人类感知本质上是多模态的,依赖文本、听觉和视觉线索来全面理解世界。为此,音视频大模型应运而生。尽管进展显著,但缺乏专用评测基准制约了模型的理解与评估。本文揭示,音视频大模型难以分辨音频与视觉信号间的细微关系,导致幻觉频发,凸显可靠评测的必要性。为此,我们提出 AVHBench,首个专为评估音视频大模型感知与理解能力设计的综合性基准。该基准涵盖幻觉检测、跨模态匹配与推理测试。实验表明,多数现有音视频大模型因对复杂多模态信号及其关系感知能力不足,在跨模态交互中幻觉率超过40%。此外,使用 AVHBench 简单训练可显著提升模型鲁棒性,减少幻觉。数据集开源:https://github.com/kaist-ami/AVHBench
原文摘要 · Abstract (English)
Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations and highlighting the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations, as well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations. Dataset: https://github.com/kaist-ami/AVHBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。