arXiv:2601.09518cs.ROcs.AI2026-01被引 10

从人类互动数据中学习人形机器人协同动作,解决接触保持与同步响应难题。

Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations

  • 提出接触导向的两阶段重定向方法,保留不同形态间的物理接触语义
  • 新模型在仿真中实现更同步、鲁棒的全身协同行为,超越单纯模仿
  • 适合研究人机协作、具身智能与动作生成的开发者与研究人员

使类人机器人实现与人类的物理交互是当前重要前沿,但高质量人-人形机器人交互(HHoI)数据稀缺。尽管人类-人类交互(HHI)数据丰富,标准重定向方法会破坏关键接触关系。为此,我们提出PAIR(物理感知交互重定向),一种以接触为中心的两阶段流程,能在形态差异下保持接触语义,生成物理一致的HHoI数据。然而,该数据暴露了传统模仿学习的缺陷:仅复制轨迹,缺乏交互理解。因此我们引入D-STAR(解耦时空动作推理器),一种分层策略,将“何时行动”与“何处行动”解耦。其中,相位注意力(when)与多尺度空间模块(where)通过扩散头融合,生成超越模仿的同步全身行为。解耦设计使模型在时间相位上更鲁棒,不受空间噪声干扰,实现响应式、协调性协作。我们在大规模严格仿真中验证框架,性能显著优于基线,构建出从HHI数据学习复杂全身交互的完整有效管道。

原文摘要 · Abstract (English)

Enabling humanoid robots to physically interact with humans is a critical frontier, but progress is hindered by the scarcity of high-quality Human-Humanoid Interaction (HHoI) data. While leveraging abundant Human-Human Interaction (HHI) data presents a scalable alternative, we first demonstrate that standard retargeting fails by breaking the essential contacts. We address this with PAIR (Physics-Aware Interaction Retargeting), a contact-centric, two-stage pipeline that preserves contact semantics across morphology differences to generate physically consistent HHoI data. This high-quality data, however, exposes a second failure: conventional imitation learning policies merely mimic trajectories and lack interactive understanding. We therefore introduce D-STAR (Decoupled Spatio-Temporal Action Reasoner), a hierarchical policy that disentangles when to act from where to act. In D-STAR, Phase Attention (when) and a Multi-Scale Spatial module (where) are fused by the diffusion head to produce synchronized whole-body behaviors beyond mimicry. By decoupling these reasoning streams, our model learns robust temporal phases without being distracted by spatial noise, leading to responsive, synchronized collaboration. We validate our framework through extensive and rigorous simulations, demonstrating significant performance gains over baseline approaches and a complete, effective pipeline for learning complex whole-body interactions from HHI data.

人机交互动作生成具身智能强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。