arXiv:2503.01921cs.CLcs.AI2025-03ACL被引 1

改进的RefChecker与SelfCheckGPT有效识别多语言文本幻觉

NCL-UoR at SemEval-2025 Task 3: Detecting Multilingual Hallucination and Related Observable Overgeneration Text Spans with Modified RefChecker and Modified SeflCheckGPT

  • 将参考文献转为基于声明的验证机制,提升事实核查精度
  • 融合外部知识增强模型自检能力,平均IoU达0.5310
  • 适合关注多语言生成质量评估的研究者使用

SemEval-2025 Task 3(Mu-SHROOM)聚焦于检测多种大型语言模型在多语言环境下生成内容中的幻觉现象。该任务不仅要求识别幻觉是否存在,还需精确定位其具体发生位置。为此,本研究提出两种改进方法:改进版RefChecker与改进版SelfCheckGPT。前者将提示式事实验证融入参考文献,将其结构化为基于声明的测试,而非单一外部知识源;后者引入外部知识以弥补其对内部知识的依赖。同时,对两种方法的原始提示设计进行优化,以识别生成文本中幻觉词汇。实验结果显示,该方法在跨语言幻觉检测任务中表现优异,在测试集上平均交并比(IoU)达0.5310,平均准确率(COR)为0.5669。

原文摘要 · Abstract (English)

SemEval-2025 Task 3 (Mu-SHROOM) focuses on detecting hallucinations in content generated by various large language models (LLMs) across multiple languages. This task involves not only identifying the presence of hallucinations but also pinpointing their specific occurrences. To tackle this challenge, this study introduces two methods: modified RefChecker and modified SelfCheckGPT. The modified RefChecker integrates prompt-based factual verification into References, structuring them as claim-based tests rather than single external knowledge sources. The modified SelfCheckGPT incorporates external knowledge to overcome its reliance on internal knowledge. In addition, both methods' original prompt designs are enhanced to identify hallucinated words within LLM-generated texts. Experimental results demonstrate the effectiveness of the approach, achieving a high ranking on the test dataset in detecting hallucinations across various languages, with an average IoU of 0.5310 and an average COR of 0.5669.

幻觉检测多语言LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。