arXiv:2512.07765cs.RO2025-12

梳理人形机器人物理交互三支柱,提出未来融合路径

Toward Seamless Physical Human-Humanoid Interaction: Insights from Control, Intent, and Modeling with a Vision for What Comes Next

  • 从建模控制、意图识别、人类模型三方面系统分析交互技术
  • 指出现有方法在动态适应与实时推断上的关键瓶颈
  • 构建交互类型分类框架,助力跨领域协同研究

物理人-人形机器人交互(pHHI)是快速发展的前沿领域,对机器人在非结构化、以人类为中心环境中的部署具有重要意义。本文从三个核心维度审视当前进展:(i) 人形机器人建模与控制,(ii) 人类意图估计,(iii) 计算人类模型。针对每个维度,综述代表性方法,识别开放挑战,并分析阻碍鲁棒、可扩展、自适应交互的现存局限,包括应对人类动态不确定性所需的全身控制策略、有限传感下的实时意图推断,以及考虑人类生理状态差异的建模技术。尽管各领域已取得显著进展,但跨领域整合仍不充分。本文提出统一方法的路径,以构建连贯的交互框架。同时,基于交互模态(直接接触与间接媒介)和机器人参与程度(辅助至协作),建立统一分类体系,为每类提供三支柱分析,明确跨域融合机会。目标是推动更鲁棒、安全、直观的物理交互,为未来研究提供路线图,使人形系统能在多样化现实场景中有效理解、预判并协同人类伙伴。

原文摘要 · Abstract (English)

Physical Human-Humanoid Interaction (pHHI) is a rapidly advancing field with significant implications for deploying robots in unstructured, human-centric environments. In this review, we examine the current state of the art in pHHI through three core pillars: (i) humanoid modeling and control, (ii) human intent estimation, and (iii) computational human models. For each pillar, we survey representative approaches, identify open challenges, and analyze current limitations that hinder robust, scalable, and adaptive interaction. These include the need for whole-body control strategies capable of handling uncertain human dynamics, real-time intent inference under limited sensing, and modeling techniques that account for variability in human physical states. Although significant progress has been made within each domain, integration across pillars remains limited. We propose pathways for unifying methods across these areas to enable cohesive interaction frameworks. This structure enables us not only to map the current landscape but also to propose concrete directions for future research that aim to bridge these domains. Additionally, we introduce a unified taxonomy of interaction types based on modality, distinguishing between direct interactions (e.g., physical contact) and indirect interactions (e.g., object-mediated), and on the level of robot engagement, ranging from assistance to cooperation and collaboration. For each category in this taxonomy, we provide the three core pillars that highlight opportunities for cross-pillar unification. Our goal is to suggest avenues to advance robust, safe, and intuitive physical interaction, providing a roadmap for future research that will allow humanoid systems to effectively understand, anticipate, and collaborate with human partners in diverse real-world settings.

人机交互人形机器人意图识别建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。