评估世界模型要聚焦决策能力,而非只是画面逼真度。
How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position
- 提出从视觉真实到策略优化的L0-L7评价阶梯框架
- 强调干预推理、闭环滚动和奖励预测等关键能力
- 适合关注机器人决策与强化学习的研学者
世界模型在现代AI中已成为核心抽象,涵盖动作条件环境模型、潜在想象模型、未来视频预测器、交互式神经模拟器、潜在预测表示和合成数据生成器等多种形式。评估标准也随之扩展,包括视频真实感、感知相似性、指令遵循、物理合理性、策略排序、可执行性、规划成功率及下游策略改进等。这导致指标多样性和主张/证据不匹配问题:部分论文对其模型的用途宣称超出评估所能证明的范围。本文综述近期文献,主张对于面向具身决策的世界模型,更关键的问题并非生成视觉逼真的视频,而是其能否支持可靠的干预推理、策略评估、规划与策略优化,尤其在干预、策略引发分布偏移和长程滚动条件下。我们构建了从视觉逼真到策略优化效用的L0–L7评价阶梯,该阶梯横跨多个正交维度,形成证据层级而非单一标量。框架突出干预动作保真度、闭环滚动有效性、奖励/价值预测、策略排序一致性、优化提升、模型可利用性及不确定性校准,并提出适用于真实机器人场景的最小报告集。
原文摘要 · Abstract (English)
World models have become a central abstraction in modern AI. The term now refers to several different objects: action-conditioned environment models, latent imagination models, future-video predictors, interactive neural simulators, latent predictive representations, and synthetic-data engines. Evaluation has broadened along with the term. Recent papers measure video realism, perceptual similarity, instruction following, physical plausibility, policy ranking, executability, planning success, and downstream policy improvement. This produces both metric diversity and a recurring problem of claim/evidence mismatch: papers sometimes make a stronger claim about what their model is useful for than their evaluation can establish. This paper surveys the recent literature and argues that, for models presented as world models for embodied decision-making, the more decisive issue is not whether the model generates visually convincing videos, but whether it supports reliable interventional reasoning, policy evaluation, planning, and policy optimization under intervention, policy-induced distribution shift, and long-horizon rollout. We organize the survey using an L0--L7 ladder spanning visual plausibility to policy optimization utility, noting that the levels cut across several orthogonal axes and so form an evidential hierarchy rather than a single scalar. The framework foregrounds interventional action fidelity, closed-loop rollout validity, reward/value prediction, policy-ranking agreement, optimization lift, model exploitability, and uncertainty calibration, with a minimal feasible reporting set for real-robot settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。