arXiv:2608.25662cs.CL2026-08综述

聚焦视觉语言模型幻觉检测,多语言细粒度识别幻觉文本片段。

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

论文配图:Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
图 1 · 摘自论文原文
  • 基于SHEEP数据集,设计跨模型幻觉检测任务。
  • 多语言下最佳系统准确率超基线30-40点。
  • 适合关注幻觉评估与生成质量的NLP研究者。

2026年,我们在与EMNLP 2026同期举办的UncertaiNLP研讨会中举办了第四届SHROOM共享任务:SHROOM-Visions(视觉语言模型中的幻觉及相关可观测过生成错误共享任务)。继2024和2025年成功举办后,本次任务旨在通过一种模型无关的检测方法,应对大视觉语言模型中的幻觉问题。任务基于新发布的SHEEP数据集,该数据集专为跨模型代际的长期评估设计,邀请参赛者在图像条件下的文本生成任务(如VQA、图像描述)中检测并分类细粒度的幻觉片段。评估采用涵盖四种语言(中文、英文、法文、意大利文)的五类幻觉分类体系。该共享任务在全球自然语言处理社区引发强烈反响,共有27支队伍提交超过600次系统评测。最佳系统在字符级相关性上达到0.58,在标签条件相关性上达0.46,在交并比(IoU)上达0.51,相比基线提升30-40个百分点。

原文摘要 · Abstract (English)

In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA, image captioning, etc.). The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian. The shared task generated strong interest in the NLP community worldwide, with 27 teams contributing 600+ system submissions. The best systems achieve average scores of 0.58 in character-level correlation, 0.46 in label-conditioned correlation, and 0.51 in intersection-over-union (IoU) across four languages, outperforming the baselines by 30-40 points.

幻觉检测多语言视觉语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。