arXiv:2604.05952cs.AIcs.CL2026-04被引 2

让论文生成更可信:通过逐步评估信心提升报告可靠性

Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration

  • 引入渐进式信心估计机制,结合深度检索与多跳推理来验证每条结论
  • 在无真实答案的开放场景中,信心评分使报告可信度提升显著
  • 适合需要高可靠性报告的研究人员与决策者使用

随着基于代理的系统不断发展,深度研究代理能够自动跨多个领域生成研究风格的报告。尽管这类代理有望简化信息整合与知识探索,但现有评估框架通常依赖主观维度,无法捕捉报告质量的关键方面——可信性。在缺乏真实答案的开放研究场景中,当前方法难以有效衡量生成内容的信念置信度,导致校准困难,使用户易受误导或幻觉信息影响。为此,我们提出一种新型深度研究代理,在报告生成流程中融入渐进式信心估计与校准机制。该系统采用反思式搜索模型,结合深度检索与多跳推理,将输出扎根于可验证证据,并为各主张分配信心评分。配合精心设计的工作流,该方法生成的报告具备更强透明性与可信度。实验结果与案例研究表明,本方法显著提升了可解释性并大幅增强用户信任。

原文摘要 · Abstract (English)

As agent-based systems continue to evolve, deep research agents are capable of automatically generating research-style reports across diverse domains. While these agents promise to streamline information synthesis and knowledge exploration, existing evaluation frameworks-typically based on subjective dimensions-fail to capture a critical aspect of report quality: trustworthiness. In open-ended research scenarios where ground-truth answers are unavailable, current evaluation methods cannot effectively measure the epistemic confidence of generated content, making calibration difficult and leaving users susceptible to misleading or hallucinated information. To address this limitation, we propose a novel deep research agent that incorporates progressive confidence estimation and calibration within the report generation pipeline. Our system leverages a deliberative search model, featuring deep retrieval and multi-hop reasoning to ground outputs in verifiable evidence while assigning confidence scores to individual claims. Combined with a carefully designed workflow, this approach produces trustworthy reports with enhanced transparency. Experimental results and case studies demonstrate that our method substantially improves interpretability and significantly increases user trust.

可信报告研究代理信心估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。