arXiv:2608.07525cs.CLcs.AI2026-08被引 1

提出自适应模糊测试框架,系统评估多模态大模型幻觉问题。

Unified Hallucination Fuzzing for Multimodal Large Language Models

论文配图:Unified Hallucination Fuzzing for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 构建细粒度统一数据集UniHall,覆盖对象、指令、知识维度。
  • 自适应模糊测试使顶尖模型幻觉率提升超60%,暴露真实可靠性短板。
  • 适合关注模型可信性与对齐风险的研究者和开发者。

幻觉仍是多模态大语言模型(MLLMs)的持续挑战,严重限制其在高风险场景中的可靠性。现有评估主要依赖静态基准,存在分类覆盖狭窄、性能快速饱和的问题,难以反映模型在动态现实场景中的鲁棒性。为此,我们提出一个整合全面基准与自演化压力测试的系统评估框架。首先,引入UniHall,一个基于统一分类体系的细粒度数据集,涵盖对象、指令、知识三个维度。其次,为应对基准饱和问题,提出自适应多模态模糊测试(SAMF),通过进化变异策略探索模型幻觉边界。关键的是,为确保动态输入下评估可靠性,SAMF采用由多模态预言机组成的结构化指标套件。大量实验表明,先进MLLMs在模糊测试下性能显著下降,与传统设置相比幻觉率提升超60%,揭示推理能力与事实一致性之间的脱节。此外,我们发现帮助性与幻觉之间存在权衡,强化学习对齐反而加剧了指令遵循任务中的阿谀倾向。框架代码与基准已开源:https://github.com/LanceZPF/EvalHall。

原文摘要 · Abstract (English)

Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing. First, we introduce UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi-modal oracles. Our extensive experiments reveal that state-of-the-art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness-hallucination trade-off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction-following tasks. The framework, code and benchmark are available at https://github.com/LanceZPF/EvalHall.

幻觉检测多模态评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。