动态评估大模型事实核查能力,自动生成测试数据并分析推理过程。
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models
- 用多智能体协作和自适应采样生成动态测试集
- 能区分顶尖大模型在事实核查中的表现差异
- 适合研究大模型可信度与推理缺陷的开发者
大语言模型(LLMs)在事实核查研究中取得显著进展。然而,现有自动化评估方法依赖静态数据集和分类指标,无法自动评估模型的推理说明生成能力,也难以揭示其在事实核查中的细微局限性。本文提出FACT-AUDIT,一种由智能体驱动的动态评估框架,可自适应、持续地评估大模型的事实核查能力。该框架基于重要性采样原理与多智能体协作,生成可扩展的动态数据集,进行迭代式以模型为中心的评估,并根据模型响应更新评价结果。通过结合判断结论与推理说明的双重评估,该框架能够全面、持续地审计大模型的事实推理能力,以探究其可信度。大量实验表明,FACT-AUDIT能有效区分当前主流大模型的表现,为以模型为中心的事实核查分析提供了宝贵洞见。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have significantly advanced the fact-checking studies. However, existing automated fact-checking evaluation methods rely on static datasets and classification metrics, which fail to automatically evaluate the justification production and uncover the nuanced limitations of LLMs in fact-checking. In this work, we introduce FACT-AUDIT, an agent-driven framework that adaptively and dynamically assesses LLMs' fact-checking capabilities. Leveraging importance sampling principles and multi-agent collaboration, FACT-AUDIT generates adaptive and scalable datasets, performs iterative model-centric evaluations, and updates assessments based on model-specific responses. By incorporating justification production alongside verdict prediction, this framework provides a comprehensive and evolving audit of LLMs' factual reasoning capabilities, to investigate their trustworthiness. Extensive experiments demonstrate that FACT-AUDIT effectively differentiates among state-of-the-art LLMs, providing valuable insights into model strengths and limitations in model-centric fact-checking analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。