arXiv:2606.18610cs.ROcs.CV2026-06被引 1

用自洽视频生成评估机器人通用策略,准确率高且可诊断失败模式。

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

论文配图:SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation
图 1 · 摘自论文原文
  • 通过动作-帧一致性、多视角一致性和测试时不确定性信号三重约束提升生成质量。
  • 在7个真实策略上达到0.929的皮尔逊相关系数和0.119的MMRV,优于已有方法。
  • 适合需要高精度评估与故障诊断的机器人策略研发团队使用。

在真实世界中评估通用机器人操作策略成本高、速度慢且难以扩展。基于动作条件的视频世界模型可通过模拟策略推演提供可扩展替代方案。自回归推演会累积误差,多视角观测需保持相互一致,且评估器须泛化至训练分布外的行为。为此,我们提出SC3-Eval,一种通过强制三种互补一致性来将预训练视频基础模型转化为精准策略评估器的自洽视频生成方法。首先,正向-逆向动力学一致性联合训练模型从动作预测帧,并从帧恢复动作,锚定生成轨迹于物理上合理的动作流形,抵消仅正向模型无法惩罚的漂移。其次,跨视角一致性训练模型从其他视角补全当前视角,确保长序列推演中多相机观测的一致性,无需显式记忆机制。第三,测试时一致性在推理阶段复用逆向动力学模块作为每动作片段的不确定性信号,终止因生成帧偏离请求动作而产生漂移的推演。我们还证明,SC3-Eval生成的轨迹能复现策略在真实推演中的失败模式,支持细粒度诊断对比而非仅聚合排名。在七个真实世界的视觉-语言-动作策略上,SC3-Eval实现闭环皮尔逊相关系数0.929和MMRV 0.119,优于三个强基线视频模型方法,并可泛化至新任务。

原文摘要 · Abstract (English)

Evaluating generalist robot manipulation policies in the real world is expensive, slow, and difficult to scale. Action-conditioned video world models offer a scalable alternative by simulating policy rollouts. Autoregressive rollouts accumulate compounding errors, observations across multiple camera views must remain mutually consistent, and the evaluator must generalize to policies whose behaviors lie outside the training distribution. We address these challenges with SC3-Eval, a self-consistent video generation recipe that adapts a pre-trained video foundation model into an accurate policy evaluator by enforcing three complementary forms of consistency. First, forward-inverse dynamics consistency jointly trains the model to predict frames from actions and to recover actions from frames, anchoring generated rollouts to a physically plausible action manifold and counteracting the drift a forward-only model cannot penalize. Second, cross-view consistency trains the model to inpaint each camera view from the other, keeping the multi-camera observation coherent over long rollouts without any explicit memory mechanism. Third, test-time consistency reuses the inverse dynamics mode at inference as a per-action-chunk uncertainty signal that terminates rollouts whose generated frames drift away from the requested actions. We also demonstrate SC3-Eval rollouts reproduce the failure modes that policies exhibit in real-world rollouts, supporting fine-grained diagnostic comparison rather than aggregate ranking alone. Across seven real-world vision-language-action policies, SC3-Eval attains a closed-loop Pearson correlation of $0.929$ and MMRV of $0.119$, outperforming three strong prior video-model-based baselines, and generalizes to new tasks.

机器人评估视频生成自洽性策略诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。