用细粒度断言审计发现大模型报告常漏关键证据,新写作框架可显著降幻觉。
Redesigning and Auditing Deep Research Writing for Faithful Reports
- 将报告拆解为断言,逐条比对来源证据
- 幻觉减少2.6至4.5倍,必要事实召回提升1.2至1.7倍
- 支持局部更新,修改源数据时重写效率更高
基于评分的深度研究(DR)系统评估常掩盖生成报告中的细微事实错误。我们提出CLAIMPROBE,一种断言级审计方法,将DR报告分解为断言,衡量幻觉、误引、引用规范性及必要事实召回率,并与检索到的证据对比。使用CLAIMPROBE发现,即使评分稳定,强效DR流水线仍可能遗漏关键证据或误引来源。随后我们提出CLAIMWRITER,一种分层断言驱动的写作框架:从源文档提取事实,映射到查询生成的提纲,并基于带源链接的断言表示撰写各部分。在三个已有DR框架中,仅替换报告生成器为CLAIMWRITER,即可使幻觉降低2.6至4.5倍,必要事实召回率提升1.2至1.7倍,同时保持整体报告质量。该方法还支持局部修订:当源文档更新时,能以最高效率将新事实融入报告,且成本更低。
原文摘要 · Abstract (English)
Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in generated reports. We introduce CLAIMPROBE, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, we find that strong DR pipelines can omit key evidence and misattribute claims even when their rubric scores remain stable. We then propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to a query-derived outline, and drafts each section from a source-linked claim representation. Across three prior DR frameworks, replacing only the report writer with CLAIMWRITER reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times, while largely preserving overall report quality. CLAIMWRITER also enables localized revision: when sources change, it propagates changed source facts into revised reports at the highest rate among update methods, while also being more cost-effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。