集体大模型决策易因微小变化产生巨大分歧,稳定性堪忧
Collective AI can amplify tiny perturbations into divergent decisions
- 多模型协作通过反复交互放大初始微小扰动
- 相同输入在12个政策场景中常导致不同最终决策
- 角色结构与记忆机制可调节系统不稳定性,适合关注AI可解释性的研究者
大型语言模型正从单体助手转向由多个成员协同讨论并投票或合成决策的委员会模式。我们发现,迭代式多模型协商反而会将微小扰动放大为迥异的对话轨迹和不同结论。在完全确定性的自托管基准测试中,重复运行结果一致,但对情景文本进行细微且语义保持不变的修改,仍会导致系统随时间分叉并改变最终建议。在部署的黑箱API系统中,即使温度设为0(本应近似确定性),多次运行仍不稳定。在12个政策场景中,这些结果表明,集体AI的不稳定性不仅源于平台侧残余随机性,更可能来自反复交互中对邻近初始条件的敏感性。额外部署实验显示,委员会架构会调节这种不稳定性:角色结构、模型构成和反馈记忆均可影响分歧程度。因此,集体AI面临的是稳定性问题,而不仅是准确性问题:确定性执行并不保证可预测或可审计的协商结果。
原文摘要 · Abstract (English)
Large language models are increasingly deployed not as single assistants but as committees whose members deliberate and then vote or synthesize a decision. Such systems are often expected to be more robust than individual models. We show that iterative multi-LLM deliberation can instead amplify tiny perturbations into divergent conversational trajectories and different final decisions. In a fully deterministic self-hosted benchmark, exact reruns are identical, yet small meaning-preserving changes to the scenario text still separate over time and often alter the final recommendation. In deployed black-box API systems, nominally identical committee runs likewise remain unstable even at temperature 0, where many users expect near-determinism. Across 12 policy scenarios, these findings indicate that instability in collective AI is not only a consequence of residual platform-side stochasticity, but can arise from sensitivity to nearby initial conditions under repeated interaction itself. Additional deployed experiments show that committee architecture modulates this instability: role structure, model composition, and feedback memory can each alter the degree of divergence. Collective AI therefore faces a stability problem, not only an accuracy problem: deterministic execution alone does not guarantee predictable or auditable deliberative outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。