arXiv:2601.03605cs.CL2026-01中稿 · EMNLP

让大模型评估事实正确性的程度,而非简单判断对错。

AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification

  • 用智能体主动搜证并精炼证据,提升判断依据质量。
  • 在多跳问答任务中,评分准确率比现有方法高12.3%。
  • 适合需要细致事实核查的AI评测与内容审核场景。

尽管大型语言模型(LLMs)取得了显著进展,其事实性仍面临严峻挑战,亟需更细致的事实验证方法。现有方法仅支持二元判断(对/错),但事实性本质上是连续谱而非非黑即白。为此,本文提出AEScorer,一种分两阶段的智能体式证据支撑框架:首先通过智能体搜索并优化外部证据,再预测一个标量事实性得分以区分细微差异。我们还构建了GradedVeriBench基准,涵盖通用与多跳问答任务。在该基准上的实验表明,AEScorer在两种设置下均显著优于现有方法,证明了针对性证据获取与分级评分结合的有效性。

原文摘要 · Abstract (English)

Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verification. Existing factuality verification methods do not capture graded judgments, even though factuality is better understood as a spectrum rather than a binary of right and wrong. To bridge this gap, we focus on graded factuality verification and propose AEScorer, an agentic evidence-grounded framework with two stages: agentic evidence acquisition and graded scoring. AEScorer first gathers and refines external evidence through agentic search, and then predicts a scalar factuality score to distinguish nuanced differences in factual correctness. We further construct GradedVeriBench, a benchmark for graded factuality verification spanning both general and multi-hop question answering. Experimental results on GradedVeriBench show that AEScorer substantially outperforms existing methods across both settings, demonstrating the value of coupling targeted evidence acquisition with graded scoring.

事实核查大模型评估智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。