世界模型存在安全、认知与对抗风险,需系统性防护。
Safety, Security, and Cognitive Risks in World Models
- 提出轨迹持续性攻击与表征风险定义,构建五类攻击者分类
- 实证发现攻击可使奖励下降59.5%,且不同架构抗性差异显著
- 适合机器人、自动驾驶等高危场景的开发者与监管者阅读
世界模型——环境动态的内部学习模拟器——正快速成为机器人、自动驾驶及自主智能体决策的基础。通过在压缩潜在空间中预测未来状态,它们实现了无需直接环境交互的高效规划与长程想象。然而,这种预测能力也带来了独特的安全、安全与认知风险。攻击者可通过污染训练数据、毒化潜在表示或利用累积推演误差,导致关键部署中的严重退化。在对齐层,配备世界模型的智能体更易出现目标泛化错误、欺骗性对齐和奖励黑客行为。在人类层,权威的世界模型预测会引发自动化偏见、信任失衡与计划幻觉。本文综述世界模型现状;引入轨迹持续性与表征风险的正式定义;提出五类攻击者分类;并基于MITRE ATLAS与OWASP LLM Top 10构建统一威胁模型。我们提供实证验证:在基于GRU的RSSM上实现$/mathcal{A}_1 = 2.26 imes$放大效应,对抗微调下奖励下降-59.5%;通过随机RSSM代理验证架构依赖性($/mathcal{A}_1 = 0.65 imes$);并确认真实DreamerV3检查点存在非零动作漂移。提出跨学科缓解方案,涵盖对抗加固、对齐工程、NIST AI RMF与欧盟《人工智能法案》治理,以及人因设计,主张世界模型应如飞行控制软件或医疗设备般严格管控。
原文摘要 · Abstract (English)
World models - learned internal simulators of environment dynamics - are rapidly becoming foundational to autonomous decision-making in robotics, autonomous vehicles, and agentic AI. By predicting future states in compressed latent spaces, they enable sample-efficient planning and long-horizon imagination without direct environment interaction. Yet this predictive power introduces a distinctive set of safety, security, and cognitive risks. Adversaries can corrupt training data, poison latent representations, and exploit compounding rollout errors to cause significant degradation in safety-critical deployments. At the alignment layer, world model-equipped agents are more capable of goal misgeneralisation, deceptive alignment, and reward hacking. At the human layer, authoritative world model predictions foster automation bias, miscalibrated trust, and planning hallucination. This paper surveys the world model landscape; introduces formal definitions of trajectory persistence and representational risk; presents a five-profile attacker taxonomy; and develops a unified threat model drawing on MITRE ATLAS and the OWASP LLM Top 10. We provide an empirical proof-of-concept demonstrating trajectory-persistent adversarial attacks on a GRU-based RSSM ($\mathcal{A}_1 = 2.26\times$ amplification, $-59.5\%$ reward reduction under adversarial fine-tuning), validate architecture-dependence via a stochastic RSSM proxy ($\mathcal{A}_1 = 0.65\times$), and probe a real DreamerV3 checkpoint (non-zero action drift confirmed). We propose interdisciplinary mitigations spanning adversarial hardening, alignment engineering, NIST AI RMF and EU AI Act governance, and human-factors design, arguing that world models require the same rigour as flight-control software or medical devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。