构建可迭代更新的论文事实性评测框架,提升大模型科研报告可信度。
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
- 通过审计-评分机制动态修订标注,让专家持续修正错误判断。
- 在四轮迭代后,专家对可验证论断的准确率从60.8%提升至90.9%。
- 适用于评估科研类大模型的事实核查能力,适合研究者与评测团队使用。
搜索增强型大语言模型代理能生成深度研究报告(DRRs),但逐条验证其事实性仍具挑战。现有事实核查工具主要针对通用领域、事实性短语类陈述设计,缺乏测试其在DRRs中迁移能力的基准。然而,构建此类基准本身困难。我们首先发现静态专家标注基准在此场景下脆弱:在博士级专家参与的对照实验中,未辅助专家在隐藏微黄金集上的准确率仅为60.8%。为此提出审计-评分(AtS)进化评测方法:当验证器与当前基准不一致时,需提交证据;由审计员裁定争议;被接受的修订将更新基准后再进行模型评分。经四轮AtS迭代,专家微黄金集准确率升至90.9%,表明专家作为审计者远比一次性标注者可靠。我们据此构建了DeepFact-Bench——一个带版本与可审计推理链的深度研究报告事实性基准,以及DeepFact-Eval——一个文档级验证代理(含轻量分组变体),在DeepFact-Bench上优于现有验证器,并在外部事实性数据集上表现出良好迁移能力。
原文摘要 · Abstract (English)
Search-augmented LLM agents can produce deep research reports (DRRs), but verifying claim-level factuality remains challenging. Existing fact-checkers are primarily designed for general-domain, factoid-style atomic claims, and there is no benchmark to test whether such verifiers transfer to DRRs. Yet building such a benchmark is itself difficult. We first show that static expert-labeled benchmarks are brittle in this setting: in a controlled study with PhD-level specialists, unassisted experts achieve only 60.8% accuracy on a hidden micro-gold set of verifiable claims. We propose Evolving Benchmarking via Audit-then-Score (AtS), where benchmark labels and rationales are explicitly revisable: when a verifier disagrees with the current benchmark, it must submit evidence; an auditor adjudicates the dispute; and accepted revisions update the benchmark before models are scored. Across four AtS rounds, expert micro-gold accuracy rises to 90.9%, indicating experts are substantially more reliable as auditors than as one-shot labelers. We instantiate AtS as DeepFact-Bench, a versioned DRR factuality benchmark with auditable rationales, and DeepFact-Eval, a document-level verification agent (with a grouped lite variant) that outperforms existing verifiers on DeepFact-Bench and transfers well to external factuality datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。