揭示AI生成内容突变的物理机制,可预测并量化幻觉风险。
Multispin Physics of AI Tipping Points and Hallucinations
- 将AI注意力头建模为多自旋热系统,发现输出突变的临界点。
- 公式表明提示词与训练偏见直接影响幻觉发生概率。
- 适用于提升AI透明度、评估用户风险与法律责任。
生成式AI如ChatGPT的输出可能重复且带有偏见,更严重的是,在响应过程中会突然从正确转为误导或错误,用户难以察觉。仅2024年,此类问题已导致670亿美元损失及数起死亡事件。通过建立与多自旋热系统的数学映射,我们揭示了在AI基本单元(注意力头)尺度上的隐藏突变不稳定性。推导出一个简洁但本质精确的突变点公式,明确展示了用户提示选择与AI训练偏见的影响。进一步说明,AI的多层架构会放大这种突变。研究成果不仅有助于提升AI透明性、可解释性与性能,还为量化用户使用AI的风险及法律责任提供了新路径。
原文摘要 · Abstract (English)
Output from generative AI such as ChatGPT, can be repetitive and biased. But more worrying is that this output can mysteriously tip mid-response from good (correct) to bad (misleading or wrong) without the user noticing. In 2024 alone, this reportedly caused $67 billion in losses and several deaths. Establishing a mathematical mapping to a multispin thermal system, we reveal a hidden tipping instability at the scale of the AI's 'atom' (basic Attention head). We derive a simple but essentially exact formula for this tipping point which shows directly the impact of a user's prompt choice and the AI's training bias. We then show how the output tipping can get amplified by the AI's multilayer architecture. As well as helping improve AI transparency, explainability and performance, our results open a path to quantifying users' AI risk and legal liabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。