arXiv:2608.08389cs.AIcs.IR2026-08

通过估算信息边际价值,减少研究型AI的冗余文本,提升效率。

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

论文配图:Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
图 1 · 摘自论文原文
  • 在检索、聚合、合成各阶段评估剪枝策略,识别最优删减时机。
  • 轻量启发式方法可节省73%令牌,几乎不降低生成质量。
  • 适合追求高效长时序推理系统的研发者参考。

长时序研究智能体通过迭代检索、聚合与综合解决开放任务,但上下文快速膨胀,而新增证据的边际价值常迅速下降,导致不必要的令牌开销、更高延迟及更嘈杂的最终报告输入。本文研究深度研究智能体中的上下文管理边际价值估计,首次系统性地对比了不同阶段的剪枝策略。我们在预检索、后检索和预合成阶段评估轻量级启发式规则与学习型价值模型。结果表明,剪枝效果更多取决于应用阶段而非评分规则:早期剪枝带来最大端到端收益,后期剪枝主要优化合成阶段上下文。轻量启发式方法可减少高达73%的令牌使用量,且质量损失极小;学习型剪枝在特定权衡下仍具竞争力,但无单一方法在质量、效率与忠实性上全面领先。这些发现为设计高效长时序智能体系统提供了实用指导。

原文摘要 · Abstract (English)

Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation. We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline. We evaluate lightweight heuristic criteria and a learned value model at pre-retrieval, post-retrieval, and pre-synthesis stages. Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness. These findings provide practical guidance for designing efficient long-horizon agentic systems.

智能体上下文压缩效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。