arXiv:2506.15107cs.ROcs.HC2025-06被引 1

为语言教学机器人定制可自适应的语音,提升表达力与易懂性。

I Know You're Listening: Adaptive Voice for HRI

  • 用表情符号引导微调Matcha-TTS,实现轻量实时的生动语音。
  • 在嘈杂环境中提高音高和语速,使语音更贴合环境感知。
  • 针对二语学习者优化元音时长,显著提升发音清晰度与理解率。

尽管社交机器人用于语言教学已受到关注,但针对第二语言教学机器人的特定合成语音研究仍有限。鉴于语言是口头任务,这一空白可能严重影响教学效果。本文提出三项贡献:1. 基于微调的Matcha-TTS,采用表情符号提示生成具时序表达力的轻量语音,支持实时运行,案例研究表明其更具表现力、社交适宜性,并适合长时间叙事。2. 探索语音随物理与社会环境自适应的方法,在嘈杂高能环境中提升音高与语速,使语音显得更契合环境且具情境感知。3. 针对二语学习者设计改进版英语语音合成系统,利用感知驱动的数据分析,发现元音时长对二语者识别难度音素对(如长/短元音)影响显著,因此构建“二语清晰模式”,仅延长长元音而保持短元音不变。该模式在测试中被证实比基线更易懂、更鼓励学习者,且显著降低转录错误率。

原文摘要 · Abstract (English)

While the use of social robots for language teaching has been explored, there remains limited work on a task-specific synthesized voices for language teaching robots. Given that language is a verbal task, this gap may have severe consequences for the effectiveness of robots for language teaching tasks. We address this lack of L2 teaching robot voices through three contributions: 1. We address the need for a lightweight and expressive robot voice. Using a fine-tuned version of Matcha-TTS, we use emoji prompting to create an expressive voice that shows a range of expressivity over time. The voice can run in real time with limited compute resources. Through case studies, we found this voice more expressive, socially appropriate, and suitable for long periods of expressive speech, such as storytelling. 2. We explore how to adapt a robot's voice to physical and social ambient environments to deploy our voices in various locations. We found that increasing pitch and pitch rate in noisy and high-energy environments makes the robot's voice appear more appropriate and makes it seem more aware of its current environment. 3. We create an English TTS system with improved clarity for L2 listeners using known linguistic properties of vowels that are difficult for these listeners. We used a data-driven, perception-based approach to understand how L2 speakers use duration cues to interpret challenging words with minimal tense (long) and lax (short) vowels in English. We found that the duration of vowels strongly influences the perception for L2 listeners and created an "L2 clarity mode" for Matcha-TTS that applies a lengthening to tense vowels while leaving lax vowels unchanged. Our clarity mode was found to be more respectful, intelligible, and encouraging than base Matcha-TTS while reducing transcription errors in these challenging tense/lax minimal pairs.

语音合成人机交互二语教学自适应语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。