研究多智能体大模型团队中领导力何时有效,发现其价值取决于任务条件。
Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams
- 用三种经典领导风格设计显式动作控制器,通过行为信号评估效果。
- 仅在初始共识不可靠时,情境型控制比随机规则提升8个百分点。
- 验证了领导力的权变理论,适合关注协作机制的研究者参考。
团队科学认为领导力是权变的:只在特定条件下有效,能力强的自主团队可能无需领导。我们探讨多智能体大模型团队中的类似问题:过程级协调控制在何种可测量条件下有价值?采用行为信号(多数锁定、探索、从错误初始共识中恢复)和逐动作消融分析,因每个控制器均为显式动作集而非单一提示。将三种经典领导风格(交易型、变革型、情境型)操作化为对共享动作词汇(探索、修订、接受、合成)的控制。与相同动作但任意规则的对照组相比,理论推导的规则能更好恢复,说明是规则而非词汇起作用。在四个任务场景和三个开源模型族中,无控制器在准确率上全面领先,符合权变观点:交易型控制在全部12个(模型,场景)组合中与初始多数投票误差小于1.3个百分点,仅在初始多数不可靠时(llama-4-scout社交任务)表现更优(情境型+8个百分点)。通过四个边界探测验证‘恢复优势’假设:控制器优于普通交互仅当初始多数不可靠、任务可修复且自由交互未自动修复时。这些区域对应权变理论中的替代、路径目标冗余、情境准备缺口,因此看似无显著提升的结果正是理论预测,非控制器失效。过程级协调控制应视为可度量、可理论映射的权变策略,而非排行榜上的胜者。
原文摘要 · Abstract (English)
Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at all. We ask the analogous question for multi-agent LLM teams: under what measurable conditions does process-level coordination control add value, and do those conditions match what team science predicts? We use behavioral signatures (majority lock-in, exploration, recovery from an incorrect round-0 consensus) and per-action ablations, clean because each controller is an explicit action set, not a monolithic prompt. We operationalize three classical leadership styles (transactional, transformational, situational) as controllers over a shared action vocabulary (explore, revise, accept, synthesize). A matched controller with the same actions but an arbitrary rule recovers no better than majority voting, so the theory-derived rule, not the vocabulary, does the work. Across four task regimes and three open-weight model families, no controller dominates by accuracy, as the contingency view predicts: transactional control matches a shared round-0 vote on all 12 (model, regime) combinations to within 1.3pp, and gains appear only on the one combination where the round-0 majority is unreliable (llama-4-scout social; situational +8pp over flat). A recovery-advantage account, tested with four boundary probes, says a controller beats plain interaction only where the round-0 majority is unreliable, the task is recoverable, and undirected interaction does not already repair it. These regions map onto contingency theory (leadership substitutes, path-goal redundancy, the situational readiness gap), so a largely null accuracy result is what the theory predicts, not a failure of the controllers. We read process-level coordination control as a contingency to be measured and theory-mapped, not a leaderboard to be topped.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。