提出细粒度多模态幻觉检测与修正方法,提升大模型输出可信度。
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- 基于细粒度跨度识别与六类错误分类,精准定位幻觉内容。
- 在6.9k人工标注样本上表现优于GPT-4o和Llama-3.2。
- 适用于需要高精度内容验证的多模态应用开发与评估。
多模态大语言模型(MLLMs)常生成与视觉内容不符的幻觉。由于幻觉形式多样,实现细粒度检测对全面评估至关重要。为此,我们提出多模态细粒度幻觉检测与编辑新任务,并提出ZINA方法:可识别细粒度幻觉片段,将其分类为六类错误类型,并建议修正方案。为训练与评估该任务,我们构建VisionHall数据集,包含12个MLLM生成的6.9k条输出,由211名标注者人工标注,以及20k条基于图结构生成的合成样本,捕捉错误类型间依赖关系。实验表明,ZINA在检测与编辑任务中均超越现有方法,包括GPT-4o与Llama-3.2。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) often generate hallucinations, where the output deviates from the visual content. Given that these hallucinations can take diverse forms, detecting hallucinations at a fine-grained level is essential for comprehensive evaluation and analysis. To this end, we propose a novel task of multimodal fine-grained hallucination detection and editing for MLLMs. Moreover, we propose ZINA, a novel method that identifies hallucinated spans at a fine-grained level, classifies their error types into six categories, and suggests appropriate refinements. To train and evaluate models for this task, we construct VisionHall, a dataset comprising 6.9k outputs from twelve MLLMs manually annotated by 211 annotators, and 20k synthetic samples generated using a graph-based method that captures dependencies among error types. We demonstrated that ZINA outperformed existing methods, including GPT-4o and Llama-3.2, in both detection and editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。