让机器人在真实环境中自然完成搬物、坐卧等交互任务
PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
- 用对抗性运动先验训练策略,模拟多样化交互行为
- 实测四类任务成功率高,跨场景泛化能力强
- 融合激光雷达与摄像头实现稳定环境感知,适合真实部署
将人形机器人部署到真实环境(如搬运物体或坐在椅子上)需要具备泛化性强、逼真的运动能力与鲁棒的场景感知。尽管以往方法分别提升了单一能力,但将其整合为统一系统仍是挑战。本文提出物理世界人形-场景交互系统 PhysHSI,使机器人能自主完成多样交互任务并保持自然动作。PhysHSI 包含仿真训练流程与真实部署系统:仿真中采用基于对抗性运动先验的策略学习,模仿多种场景下的自然交互数据,实现泛化与逼真行为;真实部署中引入粗粒度到细粒度的对象定位模块,融合 LiDAR 与相机输入,提供持续稳定的场景感知。我们在仿真与真实世界中验证了四类典型任务——搬箱、坐下、躺下、起身——均表现出高成功率、强跨任务泛化性及自然运动模式。
原文摘要 · Abstract (English)
Deploying humanoid robots to interact with real-world environments--such as carrying objects or sitting on chairs--requires generalizable, lifelike motions and robust scene perception. Although prior approaches have advanced each capability individually, combining them in a unified system is still an ongoing challenge. In this work, we present a physical-world humanoid-scene interaction system, PhysHSI, that enables humanoids to autonomously perform diverse interaction tasks while maintaining natural and lifelike behaviors. PhysHSI comprises a simulation training pipeline and a real-world deployment system. In simulation, we adopt adversarial motion prior-based policy learning to imitate natural humanoid-scene interaction data across diverse scenarios, achieving both generalization and lifelike behaviors. For real-world deployment, we introduce a coarse-to-fine object localization module that combines LiDAR and camera inputs to provide continuous and robust scene perception. We validate PhysHSI on four representative interactive tasks--box carrying, sitting, lying, and standing up--in both simulation and real-world settings, demonstrating consistently high success rates, strong generalization across diverse task goals, and natural motion patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。