用单模型统一机器人姿态估计与动作控制,提升效率与精度。
PoseDiff: A Unified Diffusion Model Bridging Robot Pose Estimation and Video-to-Action Control
- 基于扩散模型从单张图像推断3D关键点或关节角
- 在DREAM数据集上实现最优姿态估计,实时运行
- 适合需要端到端感知-规划-控制的机器人研究者
我们提出PoseDiff,一种条件扩散模型,将机器人状态估计与控制统一于单一框架中。其核心能力是从单张RGB图像中映射出结构化的机器人状态(如3D关键点或关节角),无需多阶段流程或辅助模态。在此基础上,PoseDiff可自然扩展至视频到动作的逆动力学:通过世界模型生成的稀疏视频关键帧作为条件,利用重叠平均策略生成平滑连续的长时程动作序列。该统一设计实现了感知与控制的可扩展高效集成。在DREAM数据集上,PoseDiff在姿态估计任务中达到业界最优准确率并支持实时推理;在Libero-Object操作任务中,即使在严格离线设置下,其成功率也显著优于现有逆动力学模块。结果表明,PoseDiff为具身智能中的感知、规划与控制提供了可扩展、高精度且高效的桥梁。视频演示见项目页:https://haozhuo-zhang.github.io/PoseDiff-project-page/。
原文摘要 · Abstract (English)
We present PoseDiff, a conditional diffusion model that unifies robot state estimation and control within a single framework. At its core, PoseDiff maps raw visual observations into structured robot states-such as 3D keypoints or joint angles-from a single RGB image, eliminating the need for multi-stage pipelines or auxiliary modalities. Building upon this foundation, PoseDiff extends naturally to video-to-action inverse dynamics: by conditioning on sparse video keyframes generated by world models, it produces smooth and continuous long-horizon action sequences through an overlap-averaging strategy. This unified design enables scalable and efficient integration of perception and control. On the DREAM dataset, PoseDiff achieves state-of-the-art accuracy and real-time performance for pose estimation. On Libero-Object manipulation tasks, it substantially improves success rates over existing inverse dynamics modules, even under strict offline settings. Together, these results show that PoseDiff provides a scalable, accurate, and efficient bridge between perception, planning, and control in embodied AI. The video visualization results can be found on the project page: https://haozhuo-zhang.github.io/PoseDiff-project-page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。