arXiv:2605.23929cs.AIcs.SE2026-05

优化大模型代理流程的延迟、可靠性和成本平衡。

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

论文配图:Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
图 1 · 摘自论文原文
  • 构建包含推理与输出令牌影响的可靠性模型。
  • 提出水填式令牌分配策略,实现延迟与成本约束下的最优可靠性。
  • 适合关注大模型工作流设计的系统研究人员和工程师。

现代AI系统越来越多地依赖由多个交互代理组成的流程,其中部分由大语言模型(LLM)驱动,其余由传统计算模块构成。本文分析了在大模型赋能的代理流程中延迟、可靠性和成本之间的基本权衡。我们为LLM和非LLM代理分别建立了性能模型,捕捉计算投入与输出质量的关系,并通过参数化指数可靠性函数引入推理和输出令牌的影响。随后,在延迟和成本约束下研究了串行流程的设计。主要成果包括一种水填式令牌分配策略,以及以影子价格表征的最优工作流可靠性。

原文摘要 · Abstract (English)

Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by conventional computational modules. This paper analyzes the fundamental tradeoffs between latency, reliability, and cost in LLM-enabled agentic workflows. We introduce performance models for both LLM and non-LLM agents that capture the relationship between computational effort and output quality, incorporating the impact of reasoning and output tokens for LLM agents using a parametric exponential reliability function. Then, we study the design of sequential workflows under latency and cost constraints. Main results include a water-filling token allocation policy and characterizations of optimal workflow reliability in terms of shadow prices.

大模型工作流可靠性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。