arXiv:2507.04620cs.RO2025-07中稿 · IROS 2025被引 7

通过意图识别实现人机协作的自适应切换,提升多任务场景下机器人响应能力。

IDAGC: Adaptive Generalized Human-Robot Collaboration via Human Intent Estimation and Multimodal Policy Learning

  • 融合视觉、语言、力觉等多模态数据,用CVAE估计人类意图
  • 自动识别并切换协作模式,实现在物理交互中的动态行为调整
  • 适合需要灵活人机协同的智能制造与服务机器人场景

在人机协作(HRC)中,涵盖物理交互与远程协作,准确估计人类意图并无缝切换协作模式以调整机器人行为仍是核心挑战。为此,我们提出一种基于意图驱动的自适应广义人机协作(IDAGC)框架,利用多模态数据与人类意图估计,在多样化场景中实现跨多任务的自适应策略学习,从而自主推断协作模式并动态调整机器人动作。该框架克服了现有方法通常局限于单一协作模式且无法识别和转换不同状态的局限。其核心是一个预测模型,通过条件变分自编码器(CVAE)捕捉视觉、语言、力觉与机器人状态数据间的相互依赖关系,实现高精度人类意图识别,并自动切换协作模式。各模态分别采用专用编码器,特征通过Transformer解码器融合,高效学习多任务策略;力觉数据在物理交互中同时优化合规控制与意图估计精度。实验验证了该框架在推进人机协作全面发展的实际潜力。

原文摘要 · Abstract (English)

In Human-Robot Collaboration (HRC), which encompasses physical interaction and remote cooperation, accurate estimation of human intentions and seamless switching of collaboration modes to adjust robot behavior remain paramount challenges. To address these issues, we propose an Intent-Driven Adaptive Generalized Collaboration (IDAGC) framework that leverages multimodal data and human intent estimation to facilitate adaptive policy learning across multi-tasks in diverse scenarios, thereby facilitating autonomous inference of collaboration modes and dynamic adjustment of robotic actions. This framework overcomes the limitations of existing HRC methods, which are typically restricted to a single collaboration mode and lack the capacity to identify and transition between diverse states. Central to our framework is a predictive model that captures the interdependencies among vision, language, force, and robot state data to accurately recognize human intentions with a Conditional Variational Autoencoder (CVAE) and automatically switch collaboration modes. By employing dedicated encoders for each modality and integrating extracted features through a Transformer decoder, the framework efficiently learns multi-task policies, while force data optimizes compliance control and intent estimation accuracy during physical interactions. Experiments highlights our framework's practical potential to advance the comprehensive development of HRC.

人机协作多模态意图识别自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。