用三模态输入让机器人生成自然且多样化的手势。
Gesture Generation from Trimodal Context for Humanoid Robots
- 基于语音、语义和上下文三模态输入生成手势。
- 手势与语音内容相关,风格多样,人类评价偏好明显。
- 成果成功迁移到实体机器人,提升人机交互自然度。
自然的伴随手势是提升人机交互体验的关键。然而现有手势生成方法普遍存在不自然、与语音内容不匹配或缺乏多样说话风格的问题。为此,本文旨在复现 Yoon 等人基于三模态输入在仿真中生成自然手势的工作,并将其应用于机器人。评估中采用“运动方差”和“弗雷谢手势距离(FGD)”进行客观评价,同时招募人类参与者进行主观评估。结果表明,该方法生成的手势已成功迁移至机器人,具备多样风格且与语音内容高度相关。此外,不同手势在可接受性和风格差异上均存在显著区别。
原文摘要 · Abstract (English)
Natural co-speech gestures are essential components to improve the experience of Human-robot interaction (HRI). However, current gesture generation approaches have many limitations of not being natural, not aligning with the speech and content, or the lack of diverse speaker styles. Therefore, this work aims to repoduce the work by Yoon et,al generating natural gestures in simulation based on tri-modal inputs and apply this to a robot. During evaluation, ``motion variance'' and ``Frechet Gesture Distance (FGD)'' is employed to evaluate the performance objectively. Then, human participants were recruited to subjectively evaluate the gestures. Results show that the movements in that paper have been successfully transferred to the robot and the gestures have diverse styles and are correlated with the speech. Moreover, there is a significant likeability and style difference between different gestures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。