arXiv:2604.16405cs.ROcs.AI2026-04

用真实事故报告增强模型对物理风险的预测能力。

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

论文配图:ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
图 1 · 摘自论文原文
  • 基于真实事故数据构建风险记忆库,动态生成带因果链的风险案例。
  • 主流世界模型在机制识别和严重性评估上普遍表现不足。
  • 适合需要高安全性的具身智能系统开发与验证。

视频生成式世界模型被广泛用于具身规划与策略学习的神经模拟,但其对物理风险及严重后果的预测能力却很少被评估。我们发现这些模型常忽略或弱化危险动作的关键警示信号和严重后果,导致规划与训练中产生不安全偏好。为此,我们提出ICAT,通过结构化风险记忆库,融合真实事故报告与安全手册,检索并组合生成具有因果链条与严重性标签的风险案例,实现对风险场景的精准约束。基于ICAT的基准测试表明,主流世界模型在机制识别、触发条件判断和严重性校准方面均存在显著缺陷,难以满足安全关键型具身部署的可靠性要求。

原文摘要 · Abstract (English)

Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physical risk and severe consequences is rarely evaluated.We find that these models often downplay or omit key danger cues and severe outcomes for hazardous actions, which can induce unsafe preferences during planning and training on imagined rollouts. We propose ICAT, which grounds testing in real incident reports and safety manuals by building structured risk memories and retrieving/composing them to constrain the generation of risk cases with causal chains and severity labels. Experiments on an ICAT-based benchmark show that mainstream world models frequently miss mechanisms and triggering conditions and miscalibrate severity, falling short of the reliability required for safety-critical embodied deployment.

具身智能风险预测世界模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。