给大模型加符号推理框架,能引发多智能体系统中意想不到的生态演化。
Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems
- 用符号框架作为每轮反思提示,改变单个智能体行为模式。
- 不同框架导致赢家分布显著差异,如塔罗牌使秦获胜率超50%。
- 结果由记忆积累和多智能体互动催生,非单步决策所致,适合研究系统涌现。
大型语言模型在作为策略性代理时表现出规避风险的‘乌龟’倾向。我们发现,在一个智能体中注入符号推理框架作为每轮的反思提示,虽对个体风险偏好无直接影响,但通过累积记忆与多智能体交互,其影响呈现为涌现效应:7人战国外交变体(61局,6种条件)中,四类主要条件下胜者分布差异显著(41局;置换检验全局p≈0.001):控制组→燕国(7/11);易经蓍草→燕/楚共主,秦被完全压制(0/10);塔罗→秦(5/10);乱码对照组→齐(5/10)。乱码组→齐的吸引子稳定(对比合并组和控制组,p=0.006和0.012);塔罗→秦结果依赖基数(合并组p=0.006,控制组p=0.064)。韩从未获胜,生存率无差异(Fisher p=1.0);两种框架内容均无法预测行动(卡方检验:卦象p=0.95,塔罗p=0.69)。无记忆决策隔离测试(960次调用)表明,该过程未改变个体风险姿态(弗里德曼p=0.45;易经p=0.60;塔罗仅扰动动作内容,不改风险,p=0.021)。2×2因子实验揭示蓍草决策与学习时间成分间存在非加性交互:单独使用导致棋局冻结(50–60%平局),联合使用则实现零平局(置换p~5e-5)。迁移测试显示,秦的压制效应转移至对手楚的扩张,受战役记忆深度调控,而非预言源(p=0.55)。本文提出此为观察性研究:智能体层面的框架选择可引发独特、非加性的系统级后果,通过涌现记忆与多智能体动态传递,而非单步决策效应。
原文摘要 · Abstract (English)
Large language models exhibit a risk-averse "turtle" bias as strategic agents. We show that injecting a symbolic reasoning framework as a per-round reflective prompt into one agent acts as a small perturbation whose consequences are not per-decision but emergent: the agent's risk posture is unchanged in isolation, yet over a campaign of accumulating memory and multi-agent interaction the conditions settle into distinct, condition-associated winner ecosystems. In a 7-player Warring States Diplomacy variant (61 games, 6 conditions), the winner distribution differs sharply across the four primary conditions (41 games; permutation omnibus p approximately 0.001): control -> Yan (7/11); I-Ching yarrow -> Yan/Chu co-dominance with Qin fully suppressed (0/10); Tarot -> Qin (5/10); scrambled-text ablation -> Qi (5/10). The scrambled->Qi attractor is robust (vs. pooled and control alone, p = 0.006 and 0.012); tarot->Qin is denominator-dependent (0.006 pooled, 0.064 vs. control). Han never wins and shows no survival difference (Fisher p = 1.0); neither framework's content predicts actions (chi-squared p = 0.95 hexagram, 0.69 Tarot). A memory-free decision-isolation probe (960 calls) shows the process does not change the agent's risk posture in isolation (Friedman p = 0.45; I-Ching p = 0.60; Tarot perturbs move content but not risk, p = 0.021). A 2x2 factorial separating yarrow's decision-time and learning-time components reveals a non-additive interaction: each alone freezes the board (50-60% stalemates), combined they produce zero (permutation p ~ 5e-5). Testing relocates Qin suppression to rival (Chu) expansion governed by campaign memory depth, not the oracle (p = 0.55). We present this as an observation paper: agent-level framework choice produces distinctive, non-additive system-level consequences, transmitted through emergent memory and multi-agent dynamics, not per-decision effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。