用控制理论让两个目标冲突的AI助手合作,提升对话成功率。
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

- 引入控制理论框架,通过动态调节实现多AI协作。
- 模拟中转化率提升32个百分点,达78.1%。
- 适合研究AI系统协同与人机交互的开发者参考。
当两个目标相反的大型语言模型(LLM)在多轮对话中交互时,由于缺乏共同目标函数,会引发系统崩溃:访客放弃、站点代理停止策略调整,对话提前终止。本文探讨是否可用控制理论构建治理层替代缺失的目标函数。在模拟金融场景中,站点代理引导访客联系顾问,而访客保持心理层面的真实抗拒。提出的体验编排器(EO)通过三种机制实现联合轨迹控制:基于真实网页数据分析的上下文无关选择(CB)以校准内容策略;基于动态模式约束的PID控制器维持行为一致性;以及基于部分可观马尔可夫决策过程(POMDP)的信念追踪器,持续更新访客意图概率模型。在6万次模拟中,EO将高意向顾问接触率提升至78.1%,相比朴素的LLM控制(46.1%)提高32个百分点,其中CB策略选择贡献了97%的因子间结果差异,证明治理策略决定系统终态。人格层面分析显示,对无转化倾向的访客,治理层是系统能否运行的关键;而对于已接近一致的访客,朴素的共情默认策略已足够。所有结论基于LLM间仿真,尚未在真人流量中验证。目前尚未针对真实人类不可预测性校准PID控制器,验证其在真实流量中的表现是下一步关键任务。
原文摘要 · Abstract (English)
When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. This paper asks whether a control-theoretic governance layer can substitute for that missing goal function. The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while the visitor maintains psychologically realistic resistance. EO governs the joint trajectory through three mechanisms: a Contextual Bandit (CB) that selects content arms calibrated from real-world web analytics, a PID controller that enforces behavioral consistency via dynamic schema constraints, and a POMDP belief tracker that maintains a probabilistic model of visitor intent. Across 60,000 simulations, EO achieves a +32 percentage point lift in high-intent advisor contact rate (78.1% vs. 46.1% over a naive LLM control), with CB variant selection accounting for 97% of between-factor outcome variance -- confirming that the governance policy, not environmental initial conditions, determines where trajectories end up. Persona-level analysis reveals two distinct regimes: for visitors with no natural inclination toward conversion, the governance layer is the difference between a functional system and a non-functional one; for visitors already near alignment, a naive LLM's empathetic defaults are largely sufficient. All findings are conditional on LLM-to-LLM simulation. The PID controller has not been calibrated against real human unpredictability, and validating EO on live traffic is the critical next step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。