arXiv:2607.25408cs.AI2026-07

把上下文组装当作可调控变量,提升冻结大模型智能体的稳定性与可靠性。

Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

论文配图:Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents
图 1 · 摘自论文原文
  • 将上下文组装设为控制变量,由外部策略在线学习优化
  • 证明了控制器在有限策略变动下期望奖励不下降,具备稳定性
  • 分析控制器置信度与实际任务结果的一致性,适合追求可靠性的应用

现有研究已将控制理论应用于大模型智能体:包括工具调用控制器的李雅普诺夫稳定性验证(Prinos等,2026)、大规模离散工具空间中的稀疏策略样本复杂度边界(Majumdar,2026),以及多智能体系统的可审计反馈环分解(Nogueira和Skogestad,2026)。本文不主张引入控制理论,而是聚焦于被控制的变量。此前工作控制工具选择、智能体间通信路由或原始动作流,而本文将上下文组装——即使用哪个提示模板、多少少样本示例、多少检索内容、多少次规划/验证轮次——作为受控变量,由位于冻结模型外的上下文贝叶斯或REINFORCE策略在线学习。本文建立了形式化分解(内层冻结策略π_θ,外层上下文策略π_ϕ),依据Zhang等(2026)定义给出了在线控制器的稳定性论证(在策略变化有界条件下期望奖励非递减),并报告了控制器自身置信度与实际任务结果之间的不确定性校准分析。论文的应用部分在三个领域和两个模型提供商上部署相同控制器,发布了数据集、轨迹日志和部署方案;本文重点在于形式化框架及控制论所需的稳定性与不确定性证据。

原文摘要 · Abstract (English)

A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar, "Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $π_θ$, outer context policy $π_ϕ$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.

大模型智能体控制理论上下文组装稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。