用博弈论优化大模型代理的资源分配,省17%令牌成本且质量不降。
A Stackelberg Framework for Resource-Aware LLM Agents: Learning, Repair, and Conditional Guarantees
- 将资源管理建模为上下文博弈,控制器定目标,执行者调用资源响应。
- 实测减少17.4%令牌消耗,质量无显著下降(p=0.44)。
- 适合需高效运行大模型代理的系统设计者与资源敏感场景。
大型语言模型代理越来越多地以多轮交互形式运行,必须在有限计算预算下分配上下文、提示冗长度和工具访问。静态阈值虽简单但对异构任务和动态会话状态脆弱。本文将资源治理建模为上下文型斯塔克尔伯格博弈:控制器先承诺质量目标与成本激励,执行者随后在上下文、提示和工具使用上做出资源决策。我们学习一个条件响应模型,优化领导者策略,并通过真实API校准与经验选定动作集投影来修复所得策略。对于受限博弈,我们建立了均衡存在性、跟随者响应稳定性、安全集投影及从代理环境到真实环境的迁移性等条件保证,前提是价值误差有界。主要真实API实验包含300次评估回合。相较于保守基线,所选修复后的控制器平均降低17.4%令牌成本(Welch p=0.022),而测量质量差异无统计显著性(p=0.44)。理论结果为条件成立,实验未估计其遗憾或迁移常数;因此证据仅表明一个有前景的修复运行点,而非认证的真实系统均衡。
原文摘要 · Abstract (English)
Large language model (LLM) agents increasingly operate as multi-turn systems that must allocate context, prompt verbosity, and tool access under finite computational budgets. Static thresholds are simple, but they are brittle under heterogeneous tasks and evolving session states. We formulate resource governance as a contextual Stackelberg game: a controller commits to a quality target and a cost incentive, while an executor responds with resource actions over context, prompting, and tool usage. We learn a conditional response model, optimize a leader policy against that model, and repair the resulting policy using real-API calibration and projection onto an empirically selected action set. For the restricted game, we establish conditional guarantees for equilibrium existence, follower-response stability, safe-set projection, and transfer from a surrogate environment to the real environment under bounded value error. The primary real-API experiment comprises 300 evaluated turns. Relative to a conservative baseline, the selected repaired controller reduces mean token cost by 17.4% (Welch $p=0.022$), while the measured quality difference is not statistically significant ($p=0.44$). The theoretical results are conditional and the experiments do not estimate their regret or transfer constants; consequently, the evidence establishes a promising repaired operating point, not a certified real-system equilibrium.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。