arXiv:2601.02125cs.ROcs.AI2026-01被引 2

让机器人用表情唱出丰富情感,提升人机共情体验

SingingBot: An Avatar-Driven System for Robotic Face Singing Performance

论文配图:SingingBot: An Avatar-Driven System for Robotic Face Singing Performance
图 1 · 摘自论文原文
  • 用带人类先验的视频生成模型打造生动歌唱虚拟形象
  • 通过语义映射将表情迁移到机器人,保持口型与歌声同步
  • 提出情绪动态范围指标,量化评估表演的情感丰富度

赋予机器人面部歌唱能力对实现富有同理心的人机交互至关重要。然而,现有机器人面部驱动研究主要聚焦于对话或静态表情模仿,难以满足歌唱中连续情感表达与连贯性的高要求。为此,我们提出一种新型的面向吸引力的机器人歌唱驱动框架。首先,利用嵌入广泛人类先验的肖像视频生成模型,合成生动的歌唱虚拟形象,为表情与情感提供可靠引导;随后,通过跨广泛表情空间的语义导向映射函数,将这些面部特征迁移至机器人。此外,为定量评估机器人歌唱的情感丰富度,我们提出情绪动态范围(Emotion Dynamic Range)指标,用于衡量在效价-唤醒度空间中的情感跨度,揭示广阔的情绪谱是吸引人表演的关键。全面实验表明,该方法在保持唇音同步的同时,显著提升了情感表达丰富性,优于现有方法。

原文摘要 · Abstract (English)

Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversations or mimicking static expressions, struggling to meet the high demands for continuous emotional expression and coherence in singing. To address this, we propose a novel avatar-driven framework for appealing robotic singing. We first leverage portrait video generation models embedded with extensive human priors to synthesize vivid singing avatars, providing reliable expression and emotion guidance. Subsequently, these facial features are transferred to the robot via semantic-oriented mapping functions that span a wide expression space. Furthermore, to quantitatively evaluate the emotional richness of robotic singing, we propose the Emotion Dynamic Range metric to measure the emotional breadth within the Valence-Arousal space, revealing that a broad emotional spectrum is crucial for appealing performances. Comprehensive experiments prove that our method achieves rich emotional expressions while maintaining lip-audio synchronization, significantly outperforming existing approaches.

机器人歌唱情感表达虚拟形象人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。