arXiv:2506.15085cs.ROcs.HC2025-06中稿 · RO-MAN 2025, Demo …被引 1

让机器人说话像人一样长期保持情感变化,用表情符号精细控制语调。

EmojiVoice: Towards long-term controllable expressivity in robot speech

  • 用表情符号作为提示,实现语音情感的逐段精细调节。
  • 在讲故事任务中,动态表情提示显著提升语音表现力和听众感知。
  • 轻量级模型支持机器人本地实时生成,适合部署于实际场景。

人类在长时间交谈中会不断调整语气以维持互动吸引力,而社交机器人通常使用固定欢快的语音,缺乏这种长期情感变化。虽然基础模型语音合成系统开始模仿人类语音的表达性,但难以在机器人上离线部署。我们提出 EmojiVoice,一个免费可定制的文本转语音工具包,使社交机器人能生成随时间变化的富有表现力的语音。通过引入表情符号提示,实现对语音情感的逐阶段精细控制,并采用轻量级 Matcha-TTS 架构实现实时语音生成。我们开展了三个案例研究:(1) 机器人助手的剧本对话,(2) 故事讲述机器人,(3) 语音到语音的交互代理。结果表明,在讲故事任务中,使用多样化的表情符号提示显著提升了语音感知表现力;但在助手场景中,表达性语音并不受青睐。

原文摘要 · Abstract (English)

Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with ``expressive'' joyful voices, they lack this long-term variation found in human speech. Foundation model text-to-speech systems are beginning to mimic the expressivity in human speech, but they are difficult to deploy offline on robots. We present EmojiVoice, a free, customizable text-to-speech (TTS) toolkit that allows social roboticists to build temporally variable, expressive speech on social robots. We introduce emoji-prompting to allow fine-grained control of expressivity on a phase level and use the lightweight Matcha-TTS backbone to generate speech in real-time. We explore three case studies: (1) a scripted conversation with a robot assistant, (2) a storytelling robot, and (3) an autonomous speech-to-speech interactive agent. We found that using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task, but expressive voice was not preferred in the assistant use case.

语音合成机器人交互情感控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。