用扩散模型生成真实交互场景,指导人形机器人长时序运动操作规划。
Physically Consistent Humanoid Loco-Manipulation using Latent Diffusion Models
- 通过扩散模型生成逼真图像,提取接触点与机器人构型
- 基于提取信息在全身体态优化中生成物理一致的长时序轨迹
- 适用于需要复杂交互与长期推理的人形机器人任务
本文利用潜在扩散模型(LDM)生成逼真的人机交互视觉场景,以指导人形机器人运动-操作规划。我们从生成图像中提取接触位置和机器人构型,并将其融入全身体态轨迹优化(TO)框架,生成物理上一致的运动轨迹。我们在多种长时序运动-操作场景中验证了完整流程的可行性,并对所提出的接触点与构型提取方法进行了全面分析。结果表明,利用LDM提取的信息可生成需长时序推理的物理一致性轨迹,有效支持复杂交互任务。
原文摘要 · Abstract (English)
This paper uses the capabilities of latent diffusion models (LDMs) to generate realistic RGB human-object interaction scenes to guide humanoid loco-manipulation planning. To do so, we extract from the generated images both the contact locations and robot configurations that are then used inside a whole-body trajectory optimization (TO) formulation to generate physically consistent trajectories for humanoids. We validate our full pipeline in simulation for different long-horizon loco-manipulation scenarios and perform an extensive analysis of the proposed contact and robot configuration extraction pipeline. Our results show that using the information extracted from LDMs, we can generate physically consistent trajectories that require long-horizon reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。