arXiv:2605.27073cs.LG2026-05被引 1

让智能体在不确定中自主协作,提升任务分配效率与可靠性。

Learning to Orchestrate Agents under Uncertainty

论文配图:Learning to Orchestrate Agents under Uncertainty
图 1 · 摘自论文原文
  • 将协作决策建模为带分布正则化的多臂赌博机问题
  • 在非独立同分布场景下显著优于传统基线方法
  • 适合需要动态调度异构AI模型的高不确定性任务

异构智能体的自适应协作需在不确定且动态变化的行为下做出序列化委派决策,例如协调可靠性、成本和响应质量各异的专用AI模型。现有研究多关注性能或成本,但通常未在协作层面显式建模智能体可靠性与输出分布的不确定性。本文研究在不确定性下的异构智能体自适应协作问题,提出BOT-Orch框架,将协作重构为基于智能体输出分布与任务特定参考分布之间最优传输(OT)距离正则化的多臂赌博机问题。理论证明,在标准假设下该方法具有$$\mathcal{O}(√{T})$$的遗憾界,并能对均值相同但分布对齐程度不同的智能体产生偏好排序。实验表明,在包含异构、非i.i.d.行为的合成对抗性任务分配场景中,BOT-Orch显著优于标准赌博机和启发式基线方法。

原文摘要 · Abstract (English)

Adaptive orchestration of heterogeneous agents requires making sequential delegation decisions under uncertain and evolving agent behaviour, e.g., coordinating specialised AI models with varying reliability, cost, and response quality. While prior work on agent orchestration focuses on performance or cost, uncertainty in agent reliability and output distributions is typically not modelled explicitly at the orchestration level. In this work, we study the problem of adaptive orchestration of heterogeneous agents under uncertainty, where a meta-controller must decide when to delegate to an agent, accounting for reliability, cost, and uncertainty. We propose BOT-Orch, a lightweight framework that recasts orchestration as a bandit problem over agents, regularized by OT distances between agent output distributions and task-specific reference distributions. We show that the regularised orchestration enjoys $\mathcal{O}(\sqrt{T})$ regret under standard assumptions, and provably induces preference ordering among agents with identical mean rewards but differing distributional alignment. Empirically, we demonstrate that BOT-Orch outperforms standard bandit and heuristic baselines in synthetic but adversarial task allocation settings with heterogeneous, non-i.i.d. agent behaviour.

智能体协作不确定性建模多臂赌博机优化调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。