arXiv:2601.17287cs.RO2026-01被引 1

让机器人说话时动作同步,情绪表达更自然。

Real-Time Synchronized Interaction Framework for Emotion-Aware Humanoid Robots

  • 用大模型生成文本和可执行动作描述,确保动作符合身体限制。
  • 通过动态时间对齐技术,使语音节奏与肢体动作精准匹配。
  • 实时验证动作可行性,适合医疗、教育等需要情感互动的场景。

随着人形机器人在社交场景中日益普及,实现情绪驱动的多模态同步交互仍是重大挑战。为推动人形机器人在服务角色中的进一步应用,本文提出一种面向NAO机器人的实时同步交互框架,通过三项关键创新实现语音韵律与全身动作的协调:(1) 双通道情绪引擎,利用大语言模型(LLM)同时生成情境感知的文本响应与生物力学可行的动作描述,受限于结构化关节运动库;(2) 时长感知的动态时间规整技术,实现语音输出与运动关键帧的精确时间对齐;(3) 闭环可行性验证机制,通过实时自适应确保动作符合NAO的物理关节限制。评估显示,相较于规则系统,情绪对齐度提升21%,通过协调声调(由唤醒度驱动)与上肢运动,同时保持下肢稳定。该框架实现了无缝的传感-运动协调,推动了上下文感知社交机器人在个性化医疗、互动教育及响应式客服平台等动态应用场景中的部署。

原文摘要 · Abstract (English)

As humanoid robots increasingly introduced into social scene, achieving emotionally synchronized multimodal interaction remains a significant challenges. To facilitate the further adoption and integration of humanoid robots into service roles, we present a real-time framework for NAO robots that synchronizes speech prosody with full-body gestures through three key innovations: (1) A dual-channel emotion engine where large language model (LLM) simultaneously generates context-aware text responses and biomechanically feasible motion descriptors, constrained by a structured joint movement library; (2) Duration-aware dynamic time warping for precise temporal alignment of speech output and kinematic motion keyframes; (3) Closed-loop feasibility verification ensuring gestures adhere to NAO's physical joint limits through real-time adaptation. Evaluations show 21% higher emotional alignment compared to rule-based systems, achieved by coordinating vocal pitch (arousal-driven) with upper-limb kinematics while maintaining lower-body stability. By enabling seamless sensorimotor coordination, this framework advances the deployment of context-aware social robots in dynamic applications such as personalized healthcare, interactive education, and responsive customer service platforms.

人形机器人情绪同步语音动作对齐实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。