arXiv:2509.22205cs.RO2025-09被引 3

通过人类示范视频与未来预演,实现零样本长时序机械臂操作。

From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment

  • 从非脚本化示范视频中解析出语言引导的子任务序列。
  • 生成物理合理的视觉轨迹,显式建模物体交互与接触点。
  • 在长时序任务上性能超越现有方法20%以上,适合通用机器人系统。

在零样本设置下泛化到长时序操纵任务仍是机器人领域的核心挑战。当前基于多模态基础模型的方法虽具能力,但通常仅凭静态视觉输入无法将高层指令分解为可执行的动作序列。为此,我们提出Super-Mimic,一种分层框架,通过直接从无脚本的人类示范视频中推断程序性意图,实现零样本机器人模仿。该框架包含两个顺序模块:首先,人类意图翻译器(HIT)利用多模态推理解析输入视频,生成语言接地的子任务序列;随后,未来动力学预测器(FDP)使用生成模型为每一步合成物理合理视频滚动。生成的视觉轨迹具备动力学感知,显式建模关键物体交互与接触点,以指导底层控制器。我们在一系列长时序操纵任务上进行大量实验验证,结果表明Super-Mimic显著优于当前最先进的零样本方法,性能提升超20%。这些结果表明,将视频驱动的意图解析与前瞻性动力学建模结合,是构建通用机器人系统的一种高效策略。

原文摘要 · Abstract (English)

Generalizing to long-horizon manipulation tasks in a zero-shot setting remains a central challenge in robotics. Current multimodal foundation based approaches, despite their capabilities, typically fail to decompose high-level commands into executable action sequences from static visual input alone. To address this challenge, we introduce Super-Mimic, a hierarchical framework that enables zero-shot robotic imitation by directly inferring procedural intent from unscripted human demonstration videos. Our framework is composed of two sequential modules. First, a Human Intent Translator (HIT) parses the input video using multimodal reasoning to produce a sequence of language-grounded subtasks. These subtasks then condition a Future Dynamics Predictor (FDP), which employs a generative model that synthesizes a physically plausible video rollout for each step. The resulting visual trajectories are dynamics-aware, explicitly modeling crucial object interactions and contact points to guide the low-level controller. We validate this approach through extensive experiments on a suite of long-horizon manipulation tasks, where Super-Mimic significantly outperforms state-of-the-art zero-shot methods by over 20%. These results establish that coupling video-driven intent parsing with prospective dynamics modeling is a highly effective strategy for developing general-purpose robotic systems.

机器人零样本长时序视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。