arXiv:2601.08271cs.AI2026-01被引 2

揭示大动作空间下智能体模型稳定性的稀疏性本质,为工具调用系统提供理论保障。

Sparsity Is Necessary: Polynomial-Time Stability for Agentic LLMs in Large Action Spaces

  • 基于块稀疏假设设计正则化学习方法,应对海量离散动作选择
  • 证明样本量需满足 T > k log M 才能准确识别有效工具,且稀疏性可降低样本需求
  • 适用于需要稳定控制的复杂工具调用场景,尤其适合强化学习与大模型协同系统

工具增强的大型语言模型系统面临一个学习理论长期忽视的控制范式:在巨大离散动作空间(如工具、API、文档)中进行序列决策,但每个任务分布仅依赖极小部分相关动作。本文将其形式化为稀疏智能体控制(SAC),其中策略在 M >> 1 个动作上具有块稀疏表示,奖励函数依赖稀疏主效应(可选稀疏协同效应)。通过椭球-1,2 正则化凸近似,我们建立了紧致的压缩感知风格结果:(i) 在 Policy-RSC 条件下,估计误差与价值次优性随 k (log M / T)^{1/2} 变化;(ii) 当 T > k log M 且满足非相干性和 beta-min 条件时,可通过原对偶见证法实现精确工具支持恢复;(iii) 任意稠密策略类均需 Ω(M) 样本,解释了仅靠提示的控制器不稳定的根源。进一步表明,在部分可观测情形下,LLM 的影响仅通过信念/表征误差 ε_b 体现,导致附加 O(ε_b) 退化,同时保持对 M 的对数依赖。扩展涵盖免调参、在线、鲁棒、组稀疏及交互感知的 SAC。

原文摘要 · Abstract (English)

Tool-augmented LLM systems expose a control regime that learning theory has largely ignored: sequential decision-making with a massive discrete action universe (tools, APIs, documents) in which only a small, unknown subset is relevant for any fixed task distribution. We formalize this setting as Sparse Agentic Control (SAC), where policies admit block-sparse representations over M >> 1 actions and rewards depend on sparse main effects and (optionally) sparse synergies. We study ell_{1,2}-regularized policy learning through a convex surrogate and establish sharp, compressed-sensing-style results: (i) estimation and value suboptimality scale as k (log M / T)^{1/2} under a Policy-RSC condition; (ii) exact tool-support recovery holds via primal-dual witness arguments when T > k log M under incoherence and beta-min; and (iii) any dense policy class requires Omega(M) samples, explaining the instability of prompt-only controllers. We further show that under partial observability, LLMs matter only through a belief/representation error epsilon_b, yielding an additive O(epsilon_b) degradation while preserving logarithmic dependence on M. Extensions cover tuning-free, online, robust, group-sparse, and interaction-aware SAC.

稀疏性智能体控制大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。