arXiv:2409.09318cs.CLcs.CV2024-09CVPR被引 10

提出动态评估多模态大模型幻觉的新方法,更真实地暴露模型漏洞。

ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models

  • 构建基于图的物体概念结构,动态生成多样样本
  • 在新样本上测试,发现模型幻觉率显著上升
  • 适合研究幻觉机制或优化模型的学者使用

幻觉仍是多模态大语言模型(MLLMs)面临的核心挑战。现有评估基准普遍静态,可能忽略数据污染风险。为此,我们提出 ODE——一种开放集、动态评估协议,用于在存在性和属性层面评估 MLLMs 的物体幻觉。ODE采用图结构表示现实世界中的物体概念、其属性及分布关联,据此提取符合多样化分布标准的概念组合,生成用于结构化查询的多种样本,以评估生成与判别任务中的幻觉现象。通过动态生成新样本、变化概念组合与分布频率,有效降低数据污染风险,扩展评估范围。该协议适用于通用与专用场景,包括数据有限的情况。实验表明,使用 ODE 生成的样本时,MLLMs 的幻觉率显著更高,揭示潜在的数据污染问题。同时,这些样本有助于分析幻觉模式并指导模型微调,为缓解幻觉提供了有效路径。

原文摘要 · Abstract (English)

Hallucination poses a persistent challenge for multimodal large language models (MLLMs). However, existing benchmarks for evaluating hallucinations are generally static, which may overlook the potential risk of data contamination. To address this issue, we propose ODE, an open-set, dynamic protocol designed to evaluate object hallucinations in MLLMs at both the existence and attribute levels. ODE employs a graph-based structure to represent real-world object concepts, their attributes, and the distributional associations between them. This structure facilitates the extraction of concept combinations based on diverse distributional criteria, generating varied samples for structured queries that evaluate hallucinations in both generative and discriminative tasks. Through the generation of new samples, dynamic concept combinations, and varied distribution frequencies, ODE mitigates the risk of data contamination and broadens the scope of evaluation. This protocol is applicable to both general and specialized scenarios, including those with limited data. Experimental results demonstrate the effectiveness of our protocol, revealing that MLLMs exhibit higher hallucination rates when evaluated with ODE-generated samples, which indicates potential data contamination. Furthermore, these generated samples aid in analyzing hallucination patterns and fine-tuning models, offering an effective approach to mitigating hallucinations in MLLMs.

幻觉评估多模态动态测试模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。