arXiv:2602.23271cs.AI2026-02被引 1

揭示深度研究智能体的随机性根源并提出降噪方法

Evaluating Stochasticity in Deep Research Agents

  • 将研究智能体建模为信息获取马尔可夫决策过程,量化随机性来源
  • 发现推理和早期阶段的随机性对输出变异影响最大,占比超70%
  • 通过结构化输出与集成查询生成,降低22%随机性且保持高质量

深度研究智能体(DRAs)在金融决策、医疗分析和科学发现等领域展现出巨大潜力。尽管研究质量有所提升,但其实际部署仍受制于严重的随机性:相同查询下多次运行会产生显著差异的结论与引用。本文首次形式化定义了DRAs中的随机性问题,将其建模为信息获取马尔可夫决策过程。提出评估框架,识别出信息获取、信息压缩和推理三个主要随机性来源。通过控制实验发现,推理和早期阶段的随机性对输出方差贡献最大。基于此,提出结构化输出与集成查询生成策略,可在DeepSearchQA数据集上将平均随机性降低22%,同时维持高研究质量。

原文摘要 · Abstract (English)

Deep Research Agents (DRAs) are promising agentic systems that gather and synthesize information to support research across domains such as financial decision-making, medical analysis, and scientific discovery. Despite recent improvements in research quality (e.g., outcome accuracy when ground truth is available), DRA system design often overlooks a critical barrier to real-world deployment: stochasticity. Under identical queries, repeated executions of DRAs can exhibit substantial variability in terms of research outcome, findings, and citations. In this paper, we formalize the study of stochasticity in DRAs by modeling them as information acquisition Markov Decision Processes. We introduce an evaluation framework that quantifies variance in the system and identify three sources of it: information acquisition, information compression, and inference. Through controlled experiments, we investigate how stochasticity from these modules across different decision steps influences the variance of DRA outputs. Our results show that reducing stochasticity can improve research output quality, with inference and early-stage stochasticity contributing the most to DRA output variance. Based on these findings, we propose strategies for mitigating stochasticity while maintaining output quality via structured output and ensemble-based query generation. Our experiments on DeepSearchQA show that our proposed mitigation methods reduce average stochasticity by 22% while maintaining high research quality.

智能体随机性评估框架深度研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。