构建细粒度幻觉评估基准,揭示大模型在细节视觉理解中的严重幻觉问题。
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
- 用高保真图像中的精细反常识编辑,评估多模态模型幻觉
- SOTA模型在细节感知上存在严重幻觉,评测指标仍可提升
- 通过可控子集分析提示工程对幻觉的影响,适合研究模型推理的学者
多模态大语言模型(MLLMs)存在幻觉问题。现有评估基准常因任务过于简单导致指标饱和,或多样性不足,无法充分评估先进模型的幻觉程度。为此,我们提出FREAK,一个面向MLLM细粒度幻觉评估的综合性多模态基准。通过高质量、逼真的图像,其中包含精细的反常识修改,FREAK创新性地评估模型在细节视觉感知上的幻觉现象。在FREAK上的大量实验表明,当前SOTA模型在细节视觉感知方面存在严重幻觉。为支持深入分析,我们构建了一个受控子集,间接评估模型对特定细节信息的感知能力。通过对主流链式思维(CoT)提示技术的系统评估,揭示了幻觉模式与模型推理过程的关键洞察。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) suffer from hallucinations. Existing hallucination evaluation benchmarks are often limited by over-simplified tasks leading to saturated metrics, or insufficient diversity that fails to adequately assess the hallucination extent in state-of-the-art multimodal models. To address this gap, we propose FREAK, a comprehensive multimodal benchmark designed for fine-grained hallucination assessment in MLLMs. Through high-quality photorealistic images featuring fine-grained counter-commonsense edits, FREAK innovatively evaluates hallucination phenomena in detailed visual perception of MLLMs. Extensive experiments on FREAK show severe hallucination issues in SOTA models regarding detailed visual perception. To enable deeper investigation, we curate a controlled subset to indirectly evaluate the model's ability to perceive target detailed information. Through systematic evaluation of prevailing Chain-of-Thought (CoT) prompting techniques within this task, we reveal critical insights regarding hallucination patterns and model reasoning processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。