多语言幻觉检测系统,通过检索增强定位生成文本中的虚构内容。
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
- 结合检索增强与细粒度事实拆分,识别多语言生成内容中的幻觉
- 在英语和捷克语中均位列第四,表现优于90%的参赛系统
- 适合关注多语言大模型安全性的研究者与开发者
本文提出HalluSearch,一种用于检测大型语言模型输出中虚构文本片段的多语言流水线。作为多语言幻觉与相关过生成错误共享任务(Mu-SHROOM)的一部分,HalluSearch融合检索增强验证与细粒度事实分割,可在14种不同语言中识别并定位幻觉。实证评估显示,HalluSearch在英语和捷克语任务中均排名第四,表现处于前10%水平。尽管基于检索的策略总体稳健,但在网络覆盖有限的语言中仍面临挑战,凸显了在多样化语言环境中实现一致幻觉检测仍需深入研究。
原文摘要 · Abstract (English)
In this paper, we present HalluSearch, a multilingual pipeline designed to detect fabricated text spans in Large Language Model (LLM) outputs. Developed as part of Mu-SHROOM, the Multilingual Shared-task on Hallucinations and Related Observable Overgeneration Mistakes, HalluSearch couples retrieval-augmented verification with fine-grained factual splitting to identify and localize hallucinations in fourteen different languages. Empirical evaluations show that HalluSearch performs competitively, placing fourth in both English (within the top ten percent) and Czech. While the system's retrieval-based strategy generally proves robust, it faces challenges in languages with limited online coverage, underscoring the need for further research to ensure consistent hallucination detection across diverse linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。