arXiv:2506.13130cs.CVcs.AI2025-06被引 5

提出细粒度多模态幻觉检测与修正方法,提升大模型输出可信度。

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

  • 基于细粒度跨度识别与六类错误分类,精准定位幻觉内容。
  • 在6.9k人工标注样本上表现优于GPT-4o和Llama-3.2。
  • 适用于需要高精度内容验证的多模态应用开发与评估。

多模态大语言模型(MLLMs)常生成与视觉内容不符的幻觉。由于幻觉形式多样,实现细粒度检测对全面评估至关重要。为此,我们提出多模态细粒度幻觉检测与编辑新任务,并提出ZINA方法:可识别细粒度幻觉片段,将其分类为六类错误类型,并建议修正方案。为训练与评估该任务,我们构建VisionHall数据集,包含12个MLLM生成的6.9k条输出,由211名标注者人工标注,以及20k条基于图结构生成的合成样本,捕捉错误类型间依赖关系。实验表明,ZINA在检测与编辑任务中均超越现有方法,包括GPT-4o与Llama-3.2。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) often generate hallucinations, where the output deviates from the visual content. Given that these hallucinations can take diverse forms, detecting hallucinations at a fine-grained level is essential for comprehensive evaluation and analysis. To this end, we propose a novel task of multimodal fine-grained hallucination detection and editing for MLLMs. Moreover, we propose ZINA, a novel method that identifies hallucinated spans at a fine-grained level, classifies their error types into six categories, and suggests appropriate refinements. To train and evaluate models for this task, we construct VisionHall, a dataset comprising 6.9k outputs from twelve MLLMs manually annotated by 211 annotators, and 20k synthetic samples generated using a graph-based method that captures dependencies among error types. We demonstrated that ZINA outperformed existing methods, including GPT-4o and Llama-3.2, in both detection and editing tasks.

幻觉检测多模态大模型编辑修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。