arXiv:2509.00361cs.RO2025-09被引 7

用视觉预测+通用位姿估计,让机器人在桌上自适应操作

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

  • 通过生成视频预测未来画面,指导机器人动作
  • 闭环迭代提升位姿估计精度,实现实时自适应操控
  • 无需任务专属数据,适合多样场景的机器人应用

在非结构化环境中进行机器人操作,需要系统具备跨任务泛化能力并保持稳定可靠。我们提出 GVF-TAPE 框架,将生成式视觉前瞻与任务无关的位姿估计相结合,实现可扩展的机器人桌面操作。该框架利用生成视频模型,从单视角RGB图像和任务描述中预测未来的RGB-D帧,生成可视化规划以引导机器人行为;再通过解耦的位姿估计模型从预测帧中提取末端执行器位姿,并由底层控制器转换为可执行指令。通过在闭环中迭代整合视频前瞻与位姿估计,GVF-TAPE 实现了实时、自适应的多样化任务操作。大量仿真与真实世界实验表明,该方法显著降低了对任务特定动作数据的依赖,具备良好泛化能力,为智能机器人系统提供了实用且可扩展的解决方案。

原文摘要 · Abstract (English)

Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAPE}, a closed-loop framework that combines generative visual foresight with task-agnostic pose estimation to enable scalable robotic manipulation. GVF-TAPE employs a generative video model to predict future RGB-D frames from a single side-view RGB image and a task description, offering visual plans that guide robot actions. A decoupled pose estimation model then extracts end-effector poses from the predicted frames, translating them into executable commands via low-level controllers. By iteratively integrating video foresight and pose estimation in a closed loop, GVF-TAPE achieves real-time, adaptive manipulation across a broad range of tasks. Extensive experiments in both simulation and real-world settings demonstrate that our approach reduces reliance on task-specific action data and generalizes effectively, providing a practical and scalable solution for intelligent robotic systems.

机器人操作视觉预测位姿估计闭环控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。