arXiv:2511.17792cs.CVcs.RO2025-11被引 1

首个评估视频世界模型语义规划能力的基准,发现现有模型差距明显。

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

  • 构建450个真实场景的基准,用SLAM轨迹作参考评估生成视频的运动一致性。
  • 最佳模型仅得0.341分,显示视觉生成与语义推理间存在显著鸿沟。
  • 小规模真实机器人数据微调可大幅提升规划性能,适合具身智能研究者。

尽管近期视频世界模型能生成高度逼真的视频,其进行语义推理和规划的能力仍不明确且缺乏量化评估。我们提出Target-Bench,首个全面评估视频世界模型语义推理、空间估计与规划能力的基准。该基准包含450个机器人采集的场景,覆盖47种语义类别,并以基于SLAM的轨迹作为运动趋势参考。通过度量尺度恢复机制,从生成视频中重建运动轨迹,支持五项互补指标的评估,重点关注目标接近能力和方向一致性。评估结果显示,当前最优的现成模型整体得分仅为0.341,揭示了现有视频世界模型在视觉真实性与语义推理之间的巨大差距。此外,我们证明在相对较小的真实世界机器人数据集上进行微调,可显著提升任务级规划性能。

原文摘要 · Abstract (English)

While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and unquantified. We introduce Target-Bench, the first benchmark that enables comprehensive evaluation of video world models' semantic reasoning, spatial estimation, and planning capabilities. Target-Bench provides 450 robot-collected scenarios spanning 47 semantic categories, with SLAM-based trajectories serving as motion tendency references. Our benchmark reconstructs motion from generated videos with a metric scale recovery mechanism, enabling the evaluation of planning performance with five complementary metrics that focus on target-approaching capability and directional consistency. Our evaluation result shows that the best off-the-shelf model achieves only a 0.341 overall score, revealing a significant gap between realistic visual generation and semantic reasoning in current video world models. Furthermore, we demonstrate that fine-tuning process on a relatively small real-world robot dataset can significantly improve task-level planning performance.

视频世界模型语义规划具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。