让新手用高阶语境轻松生成有表现力的语音,比现有工具更省力。
SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
- 用户输入脚本+高阶语境,系统自动优化语音表达
- 8人研究显示,生成效果更符合个人标准
- 适合内容创作者快速制作有情感的配音
新手内容创作者常需耗费大量时间录制富有表现力的语音用于社交媒体视频。尽管当前文本转语音(TTS)技术可在多种语言和口音下生成高度逼真的语音,但许多系统界面不够直观或过于细粒度。我们提出通过允许用户在脚本外添加高层级语境来简化TTS生成过程。所提出的‘SpeakEasy’原型系统基于用户提供的上下文信息调节输出,支持通过高层反馈进行迭代优化。该设计源于两项8人预研:一项调查内容创作者使用TTS的体验,另一项借鉴专业配音演员的有效策略。评估结果显示,使用SpeakEasy的参与者生成的语音表现更符合其个人标准,且所需努力程度与主流工业级界面相当。
原文摘要 · Abstract (English)
Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents, many struggle with unintuitive or overly granular TTS interfaces. We propose simplifying TTS generation by allowing users to specify high-level context alongside their script. Our Wizard-of-Oz system, SpeakEasy, leverages user-provided context to inform and influence TTS output, enabling iterative refinement with high-level feedback. This approach was informed by two 8-subject formative studies: one examining content creators' experiences with TTS, and the other drawing on effective strategies from voice actors. Our evaluation shows that participants using SpeakEasy were more successful in generating performances matching their personal standards, without requiring significantly more effort than leading industry interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。