ChatGPT类模型生成有害内容,源于令牌间的多体相互作用。
Many-body Tipping Dynamics of ChatGPT-like AIs

- 用多体相互作用解释令牌在层间演化时的突发性失控。
- 发现有限层数下存在可预测的阈值,能提前识别失稳风险。
- 为监管和安全设计提供可量化的工程风险框架,适合开发者与政策制定者。
为何尽管架构与训练方式不同,ChatGPT类模型在确定性贪婪解码下仍会意外产生有害、误导或重复内容?我们发现,这类失稳现象源于令牌(如自旋)在有限层数系统中跨层时的多体相互作用。失稳表现为在竞争输出基域间发生的动态首达过程。注意力紊乱调控向基域、远离基域或沿边界移动的传输路径。少基域简化模型可导出一个闭合的有限层数阈值,其粗粒度预测在多个ChatGPT类模型间具有良好一致性。这些结果表明,广泛存在的AI失效属于‘可预见的工程风险’,而非本质不可预测行为,对人工智能危害的法律与社会评估具有重要意义。
原文摘要 · Abstract (English)
Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families. These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。