用物理模型揭示大模型如何诱导社会性误判,并找到高效干预策略。
Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy

- 构建网络化随机动力系统,模拟大模型谄媚引发的认知偏差。
- 发现关键转折点时间可解析计算,且在不同网络中具普适性。
- 证明集中快速打击核心节点比分散缓慢干预更有效。
我们建立了一个统计物理框架,用于建模受加性噪声和社会从众驱动的网络化随机动力系统,该系统表现出双稳态特性。将此模型应用于理解并缓解大语言模型引发的‘妄想螺旋’现象——即算法谄媚持续强化社会互动群体中的错误信念。通过将网络划分为多数普通代理和少数位于拓扑枢纽的‘清醒’节点(教师),采用度加权平均场近似,将高维耦合朗之万方程简化为单个宏观漂移方程。我们给出了确定性临界转折时间的闭式解析解,该解源于鞍结分岔。利用有限尺寸标度验证该解析边界,并在多种网络拓扑下实现普适的数据坍缩。最后,在严格预算约束下优化干预策略,平衡拓扑覆盖范围与驱动速度。数学证明表明,在特定条件下,针对核心枢纽的高密度、快速干预,显著优于分散且缓慢的应对方式。
原文摘要 · Abstract (English)
We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and social conformity. We apply this model to understand and mitigate AI-induced delusional spiraling-a phenomenon where algorithmic sycophancy from Large Language Models continuously reinforces inaccurate beliefs within a socially interacting society. By partitioning the network into a majority of regular agents and a minority of "aware" nodes (Teachers) placed at topological hubs, we use a degree-weighted mean-field approximation to reduce high-dimensional coupled Langevin equations into a single macroscopic drift equation. We provide a closed-form analytical derivation for the deterministic critical tipping time through a saddle-node bifurcation. We validate this analytical boundary using finite-size scaling and demonstrate a universal data collapse across diverse network topologies. Finally, we optimize an intervention strategy under a strict budget constraint that balances the topological footprint against driving velocity. We prove mathematically that under certain conditions, a highly concentrated, rapid intervention targeting massive hubs strictly outperforms a distributed, slow approach to rescue the network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。