让语音合成支持多轮用户反馈,像导演指导演员一样精细调整发音风格。
Multi-interaction TTS toward professional recording reproduction
- 模拟导演与演员的互动过程,支持多轮语音风格修正。
- 在真实用户测试中,90%的修改请求可一次性准确实现。
- 适合需要精准控制语音风格的配音、有声书等专业场景。
语音导演常通过多轮反馈来精修配音演员的表现以达到理想效果,但这一迭代优化过程在文本转语音(TTS)领域长期被忽视。导致合成语音即使与用户期望风格存在偏差,也无法进行细粒度调整。为此,我们提出一种支持多步骤交互的TTS方法,通过建模用户与TTS模型之间的互动关系,模拟配音导演与演员的合作模式。实验表明,该方法结合自建数据集,能有效响应用户指令,实现符合预期的多轮风格修正,验证了其多交互能力。样本音频可访问:https://ntt-hilab-gensp.github.io/ssw13multiinteractiontts/
原文摘要 · Abstract (English)
Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in text-to-speech synthesis (TTS). As a result, fine-grained style refinement after the initial synthesis is not possible, even though the synthesized speech often deviates from the user's intended style. To address this issue, we propose a TTS method with multi-step interaction that allows users to intuitively and rapidly refine synthesized speech. Our approach models the interaction between the TTS model and its user to emulate the relationship between voice actors and voice directors. Experiments show that the proposed model with its corresponding dataset enables iterative style refinements in accordance with users' directions, thus demonstrating its multi-interaction capability. Sample audios are available: https://ntt-hilab-gensp.github.io/ssw13multiinteractiontts/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。