用多模型协同+内外知识,精准识别14种语言的问答幻觉
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation
- 多LLM并行模拟众包,每模型融合自身内外知识
- 在印地语上排名第一,7种语言进入前五
- 适合关注多语言幻觉检测与LLM协作的研究者
我们介绍了中央华中师范大学(CCNU)团队为Mu-SHROOM共享任务开发的系统,该任务聚焦于跨14种语言的问答系统中的幻觉识别。我们的方法利用多个具有不同专长的大型语言模型(LLMs),并行执行幻觉标注,有效模拟了众包标注过程。此外,每个基于LLM的标注器在标注过程中整合了与输入相关的内部和外部知识。使用开源LLM DeepSeek-V3,我们的系统在印地语数据上取得第一(#1)的成绩,并在另外七种语言中位列前五。本文还讨论了开发过程中探索的失败方案,并分享了参与该共享任务的关键经验。
原文摘要 · Abstract (English)
We present the system developed by the Central China Normal University (CCNU) team for the Mu-SHROOM shared task, which focuses on identifying hallucinations in question-answering systems across 14 different languages. Our approach leverages multiple Large Language Models (LLMs) with distinct areas of expertise, employing them in parallel to annotate hallucinations, effectively simulating a crowdsourcing annotation process. Furthermore, each LLM-based annotator integrates both internal and external knowledge related to the input during the annotation process. Using the open-source LLM DeepSeek-V3, our system achieves the top ranking (\#1) for Hindi data and secures a Top-5 position in seven other languages. In this paper, we also discuss unsuccessful approaches explored during our development process and share key insights gained from participating in this shared task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。