arXiv:2502.19209cs.CL2025-02被引 3

构建双语评测集与轻量模型,提升RAG幻觉检测能力

Bi'an: A Bilingual Benchmark and Model for Hallucination Detection in Retrieval-Augmented Generation

  • 基于双语数据集和轻量模型,实现多场景幻觉检测
  • 140亿参数模型性能超越五倍大模型,媲美闭源顶尖模型
  • 适合需要高效、低成本幻觉检测的中文场景研究者

检索增强生成(RAG)虽能有效减少大语言模型的幻觉,但仍可能产生不一致或无依据的内容。尽管‘大模型作为裁判’方法因实现简便被广泛采用,但面临两大挑战:缺乏全面评估基准和领域优化的裁判模型。为此,我们提出 extbf{Bi'an} 框架,包含一个双语评测数据集和轻量级裁判模型。数据集支持多种RAG场景的严格评估,裁判模型则在小型开源大模型基础上微调。在 Bi'anBench 上的大量实验表明,我们的 140 亿参数模型性能超过参数量五倍以上的基线模型,并可媲美当前最先进的闭源大模型。相关数据与模型将很快在 https://github.com/OpenSPG/KAG 开放。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) effectively reduces hallucinations in Large Language Models (LLMs) but can still produce inconsistent or unsupported content. Although LLM-as-a-Judge is widely used for RAG hallucination detection due to its implementation simplicity, it faces two main challenges: the absence of comprehensive evaluation benchmarks and the lack of domain-optimized judge models. To bridge these gaps, we introduce \textbf{Bi'an}, a novel framework featuring a bilingual benchmark dataset and lightweight judge models. The dataset supports rigorous evaluation across multiple RAG scenarios, while the judge models are fine-tuned from compact open-source LLMs. Extensive experimental evaluations on Bi'anBench show our 14B model outperforms baseline models with over five times larger parameter scales and rivals state-of-the-art closed-source LLMs. We will release our data and models soon at https://github.com/OpenSPG/KAG.

幻觉检测RAG双语轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。