arXiv:2604.11417cs.ROcs.AI2026-04被引 4

让机器人根据语义和情绪精准预测手势,无需语音输入。

Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech

论文配图:Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech
图 1 · 摘自论文原文
  • 用轻量级Transformer从文本和情绪推断手势位置与强度
  • 在BEAT2数据集上优于GPT-4o,且计算开销小
  • 适合实时部署于具身机器人,提升人机互动自然度

共言语手势能提升互动参与度并改善语言理解。现有数据驱动的机器人系统多生成节奏性敲击类动作,却很少融合语义强调。为此,我们提出一种轻量级Transformer模型,仅凭文本和情绪信息即可推断标志性手势的位置与强度,推理时无需音频输入。该模型在BEAT2数据集上的语义手势定位分类和强度回归任务中均优于GPT-4o,同时保持计算紧凑,适用于具身智能体的实时部署。

原文摘要 · Abstract (English)

Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that derives iconic gesture placement and intensity from text and emotion alone, requiring no audio input at inference time. The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression on the BEAT2 dataset, while remaining computationally compact and suitable for real-time deployment on embodied agents.

手势预测情感感知机器人交互轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。