arXiv:2607.07196cs.ROcs.AI2026-07中稿 · RSS 2026 Workshop …被引 3

提出世界模型可信性评估阶梯,确保模拟结果可信赖。

Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators

  • 构建从视觉真实到动作响应的四级可信性评估框架
  • 实测显示视觉质量高的模型动作鲁棒性反而差
  • 适用于自动驾驶等安全关键场景的模型验证

在机器人领域,世界模型(WMs)通过模拟动作后果来评估策略,但其结论可信度取决于模型自身。现有视频生成模型多以弗雷谢视频距离(FVD)衡量视觉真实性,却忽略动作响应正确性,尤其对训练中未见动作。传统仿真验证假设可信模拟器评估不可信策略,而生成式世界模型本身是未经验证的黑箱。因此我们主张,任何用作测试依据的世界模型必须先经认证。借鉴安全关键仿真中的验证、确认与认证(VV&A)、预期功能安全(SOTIF)及基于场景的测试标准,提出一个从L0到L4的可信性阶梯。该框架不依赖具体形态,以自动驾驶为例,应用在两个驾驶世界模型上发现:在视觉生成质量(L0)上排名更高的模型,在动作跟随能力(L1-L2)上反而更差,说明视觉真实度无法预测闭环评价所依赖的动作鲁棒性。

原文摘要 · Abstract (English)

Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world, and returning a success or safety verdict. Yet a verdict is only as trustworthy as the WM that produced it, and the WM itself needs to be certified. In video-generation WMs, fidelity metrics such as Fréchet Video Distance (FVD) reward visual realism, but ignore whether the world responds correctly to the policy's actions, including those unseen in training. Classical simulation-based validation assumes a trusted simulator evaluating an untrusted policy, whereas generative WMs are themselves unverified learned artifacts. Hence, we argue that any WM used as a test oracle must first be accredited before its verdicts can serve as evidence. Building on credibility practices from safety-critical simulation, including Verification, Validation & Accreditation (VV&A), Safety of the Intended Functionality (SOTIF), and scenario-based testing standards, we define an admissibility ladder (L0-L4) that a WM must climb before its closed-loop verdicts are accepted as assurance evidence. Our framework is embodiment-agnostic, and is instantiated in autonomous driving (AD), where assurance methods for traditional simulation are most mature. Applied to two driving WMs, the lower rungs reveal a reversal: the model that ranks higher on visual generation quality (L0) ranks lower on action-following (L1-L2), so visual fidelity does not predict the action-robustness a closed-loop verdict depends on.

世界模型可信评估自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。