用大模型集体判断+模糊匹配,提升多语言幻觉检测准确率
MSA at SemEval-2025 Task 3: High Quality Weak Labeling and LLM Ensemble Verification for Multilingual Hallucination Detection
- 主模型找幻觉片段,三模型投票验证真伪
- 阿拉伯语和巴斯克语检测排名第一,多语言平均表现优异
- 模拟人工标注流程,适合多语言幻觉检测研究者使用
本文介绍我们参加 SemEval-2025 Task 3:Mu-SHROOM(多语言幻觉与相关过生成错误共享任务)的参赛方案。该任务旨在跨多种语言检测指令微调大模型生成文本中的幻觉片段。我们的方法结合任务特异性提示工程与大模型集成验证机制:由主模型识别幻觉片段,再通过三个独立大模型基于概率投票判定其有效性,整体框架模拟了共享任务中验证与测试数据的人工标注流程。此外,采用模糊匹配优化片段对齐。最终系统在阿拉伯语和巴斯克语中排名第一,在德语、瑞典语和芬兰语中排名第二,在捷克语、波斯语和法语中排名第三。
原文摘要 · Abstract (English)
This paper describes our submission for SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared-task on Hallucinations and Related Observable Overgeneration Mistakes. The task involves detecting hallucinated spans in text generated by instruction-tuned Large Language Models (LLMs) across multiple languages. Our approach combines task-specific prompt engineering with an LLM ensemble verification mechanism, where a primary model extracts hallucination spans and three independent LLMs adjudicate their validity through probability-based voting. This framework simulates the human annotation workflow used in the shared task validation and test data. Additionally, fuzzy matching refines span alignment. Our system ranked 1st in Arabic and Basque, 2nd in German, Swedish, and Finnish, and 3rd in Czech, Farsi, and French.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。