评测机器人世界模型在动作指令下的预测可靠性,发现视觉逼真不等于动作准确。
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

- 构建分层评估框架,从物理一致性到动作跟随性逐级检验模型可靠性。
- 16000+人工标注判断显示:大模型未必更准,多数系统存在盲目乐观偏差。
- 适合关注机器人仿真可信度的开发者与研究者参考。
动作条件化世界模型正被广泛用作机器人学习的可扩展模拟器,但现有评估难以证明其在所依赖动作下的预测可靠性。现有基准多关注视觉保真度,却未检验预测结果是否物理合理、是否忠实于指令动作,以及在应失败时是否能正确预判失败。我们提出MiraBench,一个以动作条件化可靠性为核心目标的分层评估基准。该基准将可靠性分解为三个递进层级:物理一致性(无参考的物理合理性)、动作跟随保真度(预测是否遵循任务相关动作输入)、乐观偏差检测(在导致失败的动作下是否仍预测成功)。为此,我们构建了包含超过16,000个判断的人工标注语料库,覆盖多种任务、失败类型及主流世界模型。我们评估了12种代表性模型配置,涵盖向量条件化、文本条件化、开源/闭源系统及不同规模模型。结果显示:视觉保真度不能作为动作保真度的代理;模型规模增大并未稳定提升动作跟随能力;乐观偏差在当前系统中普遍存在。通过将评估重点从外观转向动作条件化可靠性,MiraBench为诊断和改进机器人世界模型作为可信模拟器的能力提供了基础。
原文摘要 · Abstract (English)
Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。