让机器人通过理解意图生成自然流畅的交互动作。
Hierarchical Intention-Aware Expressive Motion Generation for Humanoid Robots
- 用上下文学习识别意图,再用扩散模型实时生成动作。
- 在真实机器人上验证,动作自然且符合社交规范。
- 适合需要自然人机交互的智能服务机器人场景。
有效的人机交互要求机器人能实时识别人类意图并生成富有表现力、符合社交规范的动作。现有方法多依赖固定动作库或计算开销大的生成模型。本文提出一种分层框架,结合基于上下文学习(ICL)的意图感知推理与基于扩散模型的实时动作生成。系统采用结构化提示、置信度评分、备选行为及社交上下文感知,实现意图精炼与自适应响应。利用大规模动作数据集和高效的潜在空间去噪,该框架可生成多样且物理合理的手势,适用于动态人形机器人交互。在实体平台上的实验验证了方法在真实场景中的鲁棒性与社交契合度。
原文摘要 · Abstract (English)
Effective human-robot interaction requires robots to identify human intentions and generate expressive, socially appropriate motions in real-time. Existing approaches often rely on fixed motion libraries or computationally expensive generative models. We propose a hierarchical framework that combines intention-aware reasoning via in-context learning (ICL) with real-time motion generation using diffusion models. Our system introduces structured prompting with confidence scoring, fallback behaviors, and social context awareness to enable intention refinement and adaptive response. Leveraging large-scale motion datasets and efficient latent-space denoising, the framework generates diverse, physically plausible gestures suitable for dynamic humanoid interactions. Experimental validation on a physical platform demonstrates the robustness and social alignment of our method in realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。