构建机器人视频生成评测基准,揭示真实物理行为生成短板
Rethinking Video Generation Model for the Embodied World
- 设计跨五类任务四类机器人的标准化评测基准RBench
- 25个模型测试显示物理合理性普遍不足,人类评估相关性达0.96
- 推出400万条标注视频的RoVid-X数据集,支持高真实感训练
视频生成模型在具身智能领域取得显著进展,为生成包含感知、推理与动作的多样化机器人数据开辟新可能。然而,准确反映真实世界机器人交互的高质量视频合成仍具挑战,且缺乏统一评测基准限制了公平比较与进步。为此,我们提出综合性机器人评测基准RBench,涵盖五个任务领域和四种不同形态的机器人,通过可复现的子指标(如结构一致性、物理合理性、动作完整性)评估任务正确性与视觉保真度。对25个代表性模型的评估揭示其在生成物理合理机器人行为方面存在显著缺陷。该基准与人工评估的斯皮尔曼相关系数达0.96,验证其有效性。为突破高质量训练数据瓶颈,我们进一步提出四阶段数据处理流程,构建出目前最大的开源机器人视频数据集RoVid-X,包含400万条带注释视频片段,覆盖数千项任务,并附带全面的物理属性标注。这一评估与数据协同生态系统,为视频生成模型的严谨评估与规模化训练奠定坚实基础,推动具身智能向通用智能演进。
原文摘要 · Abstract (English)
Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing high-quality videos that accurately reflect real-world robotic interactions remains challenging, and the lack of a standardized benchmark limits fair comparisons and progress. To address this gap, we introduce a comprehensive robotics benchmark, RBench, designed to evaluate robot-oriented video generation across five task domains and four distinct embodiments. It assesses both task-level correctness and visual fidelity through reproducible sub-metrics, including structural consistency, physical plausibility, and action completeness. Evaluation of 25 representative models highlights significant deficiencies in generating physically realistic robot behaviors. Furthermore, the benchmark achieves a Spearman correlation coefficient of 0.96 with human evaluations, validating its effectiveness. While RBench provides the necessary lens to identify these deficiencies, achieving physical realism requires moving beyond evaluation to address the critical shortage of high-quality training data. Driven by these insights, we introduce a refined four-stage data pipeline, resulting in RoVid-X, the largest open-source robotic dataset for video generation with 4 million annotated video clips, covering thousands of tasks and enriched with comprehensive physical property annotations. Collectively, this synergistic ecosystem of evaluation and data establishes a robust foundation for rigorous assessment and scalable training of video models, accelerating the evolution of embodied AI toward general intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。