arXiv:2512.13713cs.AIcs.LG2025-12被引 3

用环形图测试大模型如何打破协作死循环,发现顶级模型能自创解法。

LoopBench: Discovering Emergent Symmetry Breaking Strategies with LLM Swarms

  • 让大模型在无法通信的环形图上分颜色,逼出分布式协调策略
  • 奇数环图下,普通模型会无限循环,O3等先进模型能跳出死局
  • 适合研究语言模型集体智能与自动算法生成的学者

大型语言模型(LLMs)正被越来越多地用作自主代理,但其在分布式系统中的协同能力仍不明确。我们提出 extbf{LoopBench},一个评估LLM在分布式对称性破缺与元认知思维方面推理能力的基准。该基准聚焦于使用有限颜色对奇数环图($C_3, C_5, C_{11}$)进行着色,其中无通信的确定性代理会陷入无限循环。通过引入策略传递机制作为一致记忆形式,我们发现尽管标准LLMs和经典启发式方法表现不佳,但高级推理模型(如O3)能够设计出逃逸死锁的策略。LoopBench为基于语言推理的涌现式分布式算法研究提供了测试平台,推动了集体智能的探索。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being utilized as autonomous agents, yet their ability to coordinate in distributed systems remains poorly understood. We introduce \textbf{LoopBench}, a benchmark to evaluate LLM reasoning in distributed symmetry breaking and meta-cognitive thinking. The benchmark focuses on coloring odd cycle graphs ($C_3, C_5, C_{11}$) with limited colors, where deterministic, non-communicating agents fail in infinite loops. A strategy passing mechanism is implemented as a form of consistent memory. We show that while standard LLMs and classical heuristics struggle, advanced reasoning models (e.g., O3) devise strategies to escape deadlocks. LoopBench allows the study of emergent distributed algorithms based on language-based reasoning, offering a testbed for collective intelligence.

大模型协同对称性破缺分布式算法语言智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。