用证据账本机制让AI写作更可信,自动识别证据不支持的论点。
Evidence-Ledger Adjudication for Claim-Evidence Traceability

- 为每条论点匹配证据包并判断支持关系,不支持的返回作者修改
- 在3个数据集上达到0.676准确率和0.601宏F1,远超基线
- 可精准识别1270条矛盾或缺失证据的论点,适合审稿与AI辅助写作
AI代理能快速撰写论点,但难以及时验证引用证据是否支持。本文提出证据账本裁决机制:将每条论点与证据包配对,判断支持关系,并将不支持、矛盾或混合证据的论点退回作者。核心是基于AVerTeC、CLIMATE-FEVER和SciFact三个数据集构建的2,335行盲测基准,预测时隐藏真实关系与证据标签,仅评分时合并。实验显示,该机制在基准上实现0.676的关系准确率和0.601的宏F1,显著优于最佳非代理基线(0.383准确率,0.303宏F1)。其能成功路由1270/1435条真实存在矛盾、缺失或混合证据的论点,同时仅误转295/900条真实支持的论点。结果表明,该机制可将异构证据包转化为可审计的论证可追溯层,助力可信AI写作。
原文摘要 · Abstract (English)
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。