多语言大模型幻觉检测,用维基检索+BERT识别常见幻觉模式。
TUM-MiKaNi at SemEval-2025 Task 3: Towards Multilingual and Knowledge-Aware Non-factual Hallucination Identification
- 分两阶段:先从维基检索事实,再用微调BERT识别幻觉模式。
- 在8种语言中进入前十,覆盖14种任务语言外的更多语种。
- 适合需要多语言幻觉检测的AI系统开发者和评测人员。
大语言模型的幻觉问题严重影响其可信度和应用推广。现有研究多集中于英文数据,忽视了模型的多语言特性。本文介绍我们对SemEval-2025 Task-3(Mu-SHROOM)的参赛方案,该任务聚焦多语言幻觉与相关过度生成错误。我们提出一个两阶段流水线:首先利用维基百科进行基于检索的事实验证,再通过微调BERT模型识别常见幻觉模式。系统在全部语言上表现优异,在八种语言中进入前10名,包括英文。此外,该系统支持超出任务涵盖的14种语言的多种语种。该多语言幻觉检测器有助于提升大模型输出质量,未来可广泛应用于跨语言AI系统优化。
原文摘要 · Abstract (English)
Hallucinations are one of the major problems of LLMs, hindering their trustworthiness and deployment to wider use cases. However, most of the research on hallucinations focuses on English data, neglecting the multilingual nature of LLMs. This paper describes our submission to the SemEval-2025 Task-3 - Mu-SHROOM, the Multilingual Shared-task on Hallucinations and Related Observable Overgeneration Mistakes. We propose a two-part pipeline that combines retrieval-based fact verification against Wikipedia with a BERT-based system fine-tuned to identify common hallucination patterns. Our system achieves competitive results across all languages, reaching top-10 results in eight languages, including English. Moreover, it supports multiple languages beyond the fourteen covered by the shared task. This multilingual hallucination identifier can help to improve LLM outputs and their usefulness in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。