arXiv:2504.11975cs.CL2025-04ACL被引 31

14种语言的LLM幻觉检测挑战,通过跨度标注识别生成错误。

SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes

  • 将幻觉检测建模为跨语言跨度标注任务
  • 43支队伍提交2618份结果,反映社区高度关注
  • 揭示语言差异与标注分歧带来的实际挑战

我们提出Mu-SHROOM共享任务,聚焦于指令微调的大语言模型(LLMs)输出中的幻觉及其他过生成错误的检测。该任务覆盖14种语言的通用大模型,将幻觉检测问题定义为跨度标注任务。共收到43支参赛团队提交的2,618份结果,反映出学术界对幻觉检测的高度关注。本文展示了各参赛系统的表现,并通过实证分析识别出影响性能的关键因素。同时指出当前主要挑战:不同语言间幻觉程度差异显著,且标注幻觉跨度时人工标注者间存在高分歧。

原文摘要 · Abstract (English)

We present the Mu-SHROOM shared task which is focused on detecting hallucinations and other overgeneration mistakes in the output of instruction-tuned large language models (LLMs). Mu-SHROOM addresses general-purpose LLMs in 14 languages, and frames the hallucination detection problem as a span-labeling task. We received 2,618 submissions from 43 participating teams employing diverse methodologies. The large number of submissions underscores the interest of the community in hallucination detection. We present the results of the participating systems and conduct an empirical analysis to identify key factors contributing to strong performance in this task. We also emphasize relevant current challenges, notably the varying degree of hallucinations across languages and the high annotator disagreement when labeling hallucination spans.

幻觉检测多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。