用物理理论解析大模型注意力机制,揭示幻觉与偏见根源
Capturing AI's Attention: Physics of Repetition, Hallucination, Bias and Beyond
- 从第一性原理构建注意力头的物理模型
- 预测输出重复、幻觉等问题并匹配大规模实验结果
- 类自旋浴结构可借力物理学提升AI可信度
我们推导出大语言模型核心‘注意力头’的第一性原理物理理论,可定量分析输出重复、幻觉、有害内容及训练/微调带来的偏见等关键挑战。理论预测与大规模大模型输出一致。其二体形式解释了大模型为何有效,暗示三体注意力可能进一步提升性能。该模型与自旋浴系统相似,表明现有物理学知识可直接用于提升AI的可信度与抗操纵能力。
原文摘要 · Abstract (English)
We derive a first-principles physics theory of the AI engine at the heart of LLMs' 'magic' (e.g. ChatGPT, Claude): the basic Attention head. The theory allows a quantitative analysis of outstanding AI challenges such as output repetition, hallucination and harmful content, and bias (e.g. from training and fine-tuning). Its predictions are consistent with large-scale LLM outputs. Its 2-body form suggests why LLMs work so well, but hints that a generalized 3-body Attention would make such AI work even better. Its similarity to a spin-bath means that existing Physics expertise could immediately be harnessed to help Society ensure AI is trustworthy and resilient to manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。