用熵动力学分析大模型多智能体系统,发现推理越强越容易卡住。
Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems

- 构建平均场熵动力学模型,模拟任务解决与上下文堆积的博弈
- 实测显示强推理模型在协调时易因上下文过载而崩溃
- 提出逆向工作流生成法,可构造带中间检查点的复杂测试集
从单轮模型转向多智能体系统(MAS)有望提升问题求解能力,但集中式协调拓扑仍是关键脆弱点。为此,我们提出均值场熵动力学框架,将协调过程建模为任务解决与累积上下文加载之间竞争作用的系统。为支持验证,我们引入逆向工作流生成(IWG),一种可合成具有密集中间检查点、过程可验证的高复杂度基准的多智能体流水线。我们证明熵动力学模型能拟合实际轨迹,提供可物理解释的参数,量化系统稳定性和性能崩溃。关键发现是存在‘推理陷阱’:尽管推理能力强的模型在独立任务中表现优异,但在作为协调者时却常因上下文压缩而失败。揭示协调者背后的物理机制并量化系统不确定性,有助于多智能体系统的架构设计。
原文摘要 · Abstract (English)
The transition from single-turn models to Multi-Agent Systems (MAS) promises enhanced problem-solving capabilities, yet the centralized orchestration topology remains a critical point of fragility. To analyze this, we propose a Mean-Field Entropy Dynamics framework, modeling the orchestration process as a system governed by the competing forces of task resolution and cumulative context loading. To facilitate validation, we introduce Inverse Workflow Generation (IWG), a multi-agent pipeline that synthesizes process-verifiable, high-complexity benchmarks with dense intermediate checkpoints. We demonstrate that our entropy dynamics model fits empirical trajectories, providing physically interpretable parameters that quantify system stability and performance collapse. Crucially, our analysis uncovers a ``Reasoning Trap": while reasoning-heavy models excel in isolated tasks, they frequently fail as orchestrators due to context squeezing. Elucidating the physical mechanisms underlying the Orchestrator and quantifying systemic uncertainty offers insights for the MASs' architectural design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。