arXiv:2603.18771cs.RO2026-03

让教育机器人根据教学内容生成有情感、有逻辑的肢体动作。

Empathetic Motion Generation for Humanoid Educational Robots via Reasoning-Guided Vision--Language--Motion Diffusion Architecture

  • 用多模态融合预测情绪,再通过教学推理生成动作类别。
  • 生成动作更结构化且与指令匹配,比基线模型提升显著。
  • 适合开发智能教育机器人,尤其关注互动表达力的研究者。

本文提出一种基于推理引导的视觉-语言-动作扩散框架(RG-VLMD),用于在教育场景中生成与指令相关的伴随动作。系统整合多模态情感估计、教学推理与教学行为条件化的运动合成,实现自适应且语义一致的机器人行为。通过门控专家混合模型从文本、视觉和音频特征中预测情感维度(效价/唤醒度),并映射为离散教学行为类别,作为扩散运动生成器的条件。该生成器采用片段级意图与帧级教学计划进行附加潜在约束,并辅以动作组监督。相较于基线扩散模型,所提方法生成的动作更具结构性与区分性,经运动统计与成对距离分析验证。生成序列保持物理合理性,可实时重定向至NAO机器人执行。结果表明,推理引导的教学条件化显著提升了动作可控性与教学表现力。

原文摘要 · Abstract (English)

This article suggests a reasoning-guided vision-language-motion diffusion framework (RG-VLMD) for generating instruction-aware co-speech gestures for humanoid robots in educational scenarios. The system integrates multi-modal affective estimation, pedagogical reasoning, and teaching-act-conditioned motion synthesis to enable adaptive and semantically consistent robot behavior. A gated mixture-of-experts model predicts Valence/Arousal from input text, visual, and acoustic features, which then mapped to discrete teaching-act categories through an affect-driven policy.These signals condition a diffusion-based motion generator using clip-level intent and frame-level instructional schedules via additive latent restriction with auxiliary action-group supervision. Compared to a baseline diffusion model, our proposed method produces more structured and distinctive motion patterns, as verified by motion statics and pairwise distance analysis. Generated motion sequences remain physically plausible and can be retargeted to a NAO robot for real-time execution. The results reveal that reasoning-guided instructional conditioning improves gesture controllability and pedagogical expressiveness in educational human-robot interaction.

教育机器人动作生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。