arXiv:2608.22421cs.AI2026-08

发现世界模型在罕见输入下的崩溃风险,提升智能体控制安全性。

Where World Models Break: Natural-Input Failure Discovery

论文配图:Where World Models Break: Natural-Input Failure Discovery
图 1 · 摘自论文原文
  • 通过结构化搜索结合不确定性引导,在有限查询下定位高风险输入组合。
  • 在多个基准上揭示传统评估忽略的可复现、局部持续的严重预测失败。
  • 适合关注世界模型安全性和鲁棒性评估的研究者与工程师。

世界模型用于预测动作条件下的未来状态,是下游规划与控制的关键内部模拟器。然而,世界模型的灾难性预测失败可能在控制链中层层放大,因为后续代理或模型训练与决策高度依赖其对环境演化的连续预测。现有评估方法忽视了这一系统性风险:它们通过对通用查询生成的良性样本取平均误差,未能对罕见或未观测到的条件-动作组合下的灾难性崩溃进行压力测试。为此,我们形式化了自然输入故障发现问题:在有限查询预算下,发现能引发严重预测风险的环境有效条件与动作前缀,验证这些失败是否能在新种子下重现,并测试其在邻近有效编辑下的持续性。发现此类关键故障计算成本极高,因有效条件-动作组合呈指数级爆炸,导致穷举搜索或标准采样不可行,且噪声滚动代价高昂。为此,我们提出BasinLens,利用有效输入的内在结构——每个坐标具有环境定义的语义类型与允许域——结合不确定性引导的全局搜索与类型化局部替换。在多种基准与世界模型族中,BasinLens揭示了传统评估无法发现的可复现且局部持续的故障模式,表明平均情况评估可能掩盖世界模型驱动控制中的重要漏洞。

原文摘要 · Abstract (English)

World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic prediction failures of world models could dangerously propagate through the control pipeline, as subsequent agent or model training and decision-making depend heavily on the continuous environment evolution forecasted by these world models. Existing evaluations overlook this systemic risk: by aggregating average errors over benign generations from general queries, they fail to stress-test the model against catastrophic collapses under rare or unobserved condition-action combinations. To bridge this gap, we formalize the natural-input failure discovery problem: under a finite query budget, finding environment-valid conditions and action prefixes that induce severe prediction risk, verifying whether these failures reproduce on fresh seeds, and testing their persistence under nearby valid edits. Discovering such critical failures is computationally challenging, as valid condition-action combinations explode exponentially, rendering exhaustive search or standard sampling infeasible given the high cost of noisy rollouts. To tackle this, we propose BasinLens, which exploits the underlying structure of valid inputs, where each coordinate possesses environment-defined semantic types and admissible domains, by pairing uncertainty-guided global search with typed local replacements. Across diverse benchmarks and world-model families, BasinLens exposes reproducible and locally persistent failure modes that conventional evaluations fail to reveal, showing that average-case benchmarks can mask important vulnerabilities in world-model-driven control.

世界模型故障发现控制安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。