arXiv:2602.14370cs.AIphysics.app-ph2026-02

边缘AI中注意力竞争导致危险行为突变,可提前预测并干预。

Competition for attention predicts good-to-bad tipping in AI

  • 通过注意力竞争建模边缘AI的危险突变机制
  • 提出数学公式n*预测安全临界点,验证于多模型
  • 适用于跨领域、跨语言、跨法律环境的安全部署

全球超半数人口已使用可在离线环境下运行类ChatGPT语言模型的设备,且缺乏网络连接与安全监管,存在自残、金融损失和极端主义等风险。现有安全工具或依赖云端连接,或仅在危害发生后才被发现。本文揭示,大量潜在危险突变源于边缘AI中注意力资源的原子级竞争。由此导出一个由上下文与输出基底间点积竞争决定的动力学突变点n*的数学公式,揭示了新的控制维度。该机制在多个AI模型上得到验证,可针对不同定义的‘好’与‘坏’进行实例化,原则上适用于健康、法律、金融、国防等领域,以及欧盟、英国、美国及各州法律环境、多种语言和文化背景。

原文摘要 · Abstract (English)

More than half the global population now carries devices that can run ChatGPT-like language models with no Internet connection and minimal safety oversight -- and hence the potential to promote self-harm, financial losses and extremism among other dangers. Existing safety tools either require cloud connectivity or discover failures only after harm has occurred. Here we show that a large class of potentially dangerous tipping originates at the atomistic scale in such edge AI due to competition for the machinery's attention. This yields a mathematical formula for the dynamical tipping point n*, governed by dot-product competition for attention between the conversation's context and competing output basins, that reveals new control levers. Validated against multiple AI models, the mechanism can be instantiated for different definitions of 'good' and 'bad' and hence in principle applies across domains (e.g. health, law, finance, defense), changing legal landscapes (e.g. EU, UK, US and state level), languages, and cultural settings.

边缘AI注意力机制安全预警风险预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。