arXiv:2507.06235cs.HCcs.AI2025-07被引 9

探索让电脑语音更可爱的声学秘诀,发现特定音高与音色可增强萌感。

Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice

  • 通过调节基频和共振峰,手动与自动操控语音特征
  • 在512人实验中发现语音萌感存在最优区间与天花板效应
  • 适合研究人机交互、虚拟角色设计及语音情感表达的开发者

“可爱”(Kawaii)是日本文化中与社会身份和情感反应相关的概念,但现有研究多聚焦视觉层面,忽视了声音中的可爱性。为建立声音可爱性的科学体系,本研究探索语音特征与可爱感的关系,并实现人工与自动操控。通过四阶段研究(总样本量512),对比文本转语音(TTS)与游戏角色语音,发现特定基频与共振峰组合可激发“可爱甜点”,但效果因语音类型而异且存在上限。研究验证了初步的声音可爱性模型,并提出一种基础操控方法,为计算机语音的情感化设计提供实证支持。

原文摘要 · Abstract (English)

"Kawaii" is the Japanese concept of cute, which carries sociocultural connotations related to social identities and emotional responses. Yet, virtually all work to date has focused on the visual side of kawaii, including in studies of computer agents and social robots. In pursuit of formalizing the new science of kawaii vocalics, we explored what elements of voice relate to kawaii and how they might be manipulated, manually and automatically. We conducted a four-phase study (grand N = 512) with two varieties of computer voices: text-to-speech (TTS) and game character voices. We found kawaii "sweet spots" through manipulation of fundamental and formant frequencies, but only for certain voices and to a certain extent. Findings also suggest a ceiling effect for the kawaii vocalics of certain voices. We offer empirical validation of the preliminary kawaii vocalics model and an elementary method for manipulating kawaii perceptions of computer voice.

语音生成情感计算人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。