arXiv:2604.08664cs.RO2026-04

用自然语言生成物理人机交互场景,实现零样本仿真到现实的迁移。

Generative Simulation for Policy Learning in Physical Human-Robot Interaction

  • 通过大模型自动生成人体、场景和机器人动作,实现零样本场景合成。
  • 在抓痒和洗澡任务中成功实现80%以上成功率,对非预设人体动作有鲁棒性。
  • 适合做智能助手机器人训练,尤其适用于缺乏真实数据的场景。

构建自主物理人机交互(pHRI)系统受限于大规模训练数据的匮乏。本文提出一种零样本「文本到仿真到现实」生成式仿真框架,可从自然语言提示自动合成多样化的pHRI场景。该框架利用大语言模型(LLMs)和视觉-语言模型(VLMs),程序化生成软体人体模型、场景布局及机器人运动轨迹,用于辅助任务。我们基于此框架自动收集大规模合成示范数据集,并训练基于分割点云的视觉模仿学习策略。通过两项物理辅助任务(抓痒与洗澡)的用户研究验证,所学策略实现零样本仿真到现实迁移,在未预设人体动作下仍保持超80%成功率。本工作首次实现pHRI应用中仿真环境合成、数据采集与策略学习的全流程自动化。

原文摘要 · Abstract (English)

Developing autonomous physical human-robot interaction (pHRI) systems is limited by the scarcity of large-scale training data to learn robust robot behaviors for real-world applications. In this paper, we introduce a zero-shot "text2sim2real" generative simulation framework that automatically synthesizes diverse pHRI scenarios from high-level natural-language prompts. Leveraging Large Language Models (LLMs) and Vision-Language Models (VLMs), our pipeline procedurally generates soft-body human models, scene layouts, and robot motion trajectories for assistive tasks. We utilize this framework to autonomously collect large-scale synthetic demonstration datasets and then train vision-based imitation learning policies operating on segmented point clouds. We evaluate our approach through a user study on two physically assistive tasks: scratching and bathing. Our learned policies successfully achieve zero-shot sim-to-real transfer, attaining success rates exceeding 80% and demonstrating resilience to unscripted human motion. Overall, we introduce the first generative simulation pipeline for pHRI applications, automating simulation environment synthesis, data collection, and policy learning. Additional information may be found on our project website: https://rchi-lab.github.io/gen_phri/

人机交互生成仿真模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。