构建中文幻觉评估基准,自动生成高质量问答数据。
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
- 用智能体自动从网页抓取知识文档构造细粒度问答数据
- 基于1399份文档生成60702条数据,覆盖16个主流大模型
- 适合需要中文幻觉检测与评估的研究者使用
尽管大语言模型发展迅速,仍极易产生幻觉,严重制约其应用。现有幻觉评估多依赖人工标注,难以实现自动化与低成本评估。为此,我们提出HaluAgent智能体框架,可基于知识文档自动生成细粒度QA数据。实验表明,手动设计规则与提示优化能有效提升数据质量。利用该框架,我们构建了C-FAITH,一个源自1,399份网络抓取文档的中文问答幻觉评估基准,共包含60,702条样本。我们对16个主流大语言模型进行了全面评估,提供了详尽的实验结果与分析。
原文摘要 · Abstract (English)
Despite the rapid advancement of large language models, they remain highly susceptible to generating hallucinations, which significantly hinders their widespread application. Hallucination research requires dynamic and fine-grained evaluation. However, most existing hallucination benchmarks (especially in Chinese language) rely on human annotations, making automatical and cost-effective hallucination evaluation challenging. To address this, we introduce HaluAgent, an agentic framework that automatically constructs fine-grained QA dataset based on some knowledge documents. Our experiments demonstrate that the manually designed rules and prompt optimization can improve the quality of generated data. Using HaluAgent, we construct C-FAITH, a Chinese QA hallucination benchmark created from 1,399 knowledge documents obtained from web scraping, totaling 60,702 entries. We comprehensively evaluate 16 mainstream LLMs with our proposed C-FAITH, providing detailed experimental results and analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。