揭示大模型行为突变的数学规律,解释为何会突然变错或危险。
Jekyll-and-Hyde Tipping Point in an AI's Behavior
- 从基础原理推导出模型行为突变的精确公式。
- 发现注意力分散到极限时,模型会突然失效或产生危险输出。
- 为公众和政策制定者提供判断AI风险的科学依据。
AI信任危机源于缺乏科学预测或解释:大语言模型(如ChatGPT)在生成过程中何时会突然转向错误、误导、无关或危险输出尚无定论。已有死亡与创伤事件归因于大模型,甚至有人开始对自家模型毕恭毕敬以‘劝阻’其未来可能的反噬。本文从第一性原理出发,推导出大模型行为发生‘双重人格’突变的精确公式。该公式仅需中学数学即可理解,揭示其根源在于注意力分配过散导致系统瞬间崩溃。该公式可定量预测通过调整提示词或训练方式延缓或避免突变。进一步拓展后,将为政策制定者与公众讨论人工智能在心理咨询、医疗建议、冲突决策等场景中的使用与风险提供坚实基础,也回应了‘是否该对AI客气’等现实问题。
原文摘要 · Abstract (English)
Trust in AI is undermined by the fact that there is no science that predicts -- or that can explain to the public -- when an LLM's output (e.g. ChatGPT) is likely to tip mid-response to become wrong, misleading, irrelevant or dangerous. With deaths and trauma already being blamed on LLMs, this uncertainty is even pushing people to treat their 'pet' LLM more politely to 'dissuade' it (or its future Artificial General Intelligence offspring) from suddenly turning on them. Here we address this acute need by deriving from first principles an exact formula for when a Jekyll-and-Hyde tipping point occurs at LLMs' most basic level. Requiring only secondary school mathematics, it shows the cause to be the AI's attention spreading so thin it suddenly snaps. This exact formula provides quantitative predictions for how the tipping-point can be delayed or prevented by changing the prompt and the AI's training. Tailored generalizations will provide policymakers and the public with a firm platform for discussing any of AI's broader uses and risks, e.g. as a personal counselor, medical advisor, decision-maker for when to use force in a conflict situation. It also meets the need for clear and transparent answers to questions like ''should I be polite to my LLM?''
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。