让AI用辩论风格生成反驳语音,无需训练新声音
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
- 用对手发言做语调参考,自己声音做身份标识
- 在真实数据上预训练,支持任意声音的辩论式语音合成
- 适合需要多角色辩论语音生成的应用场景
辩论中,反驳是关键环节,发言者需根据对方论点生成有说服力的回应。本文提出一种零样本文本到语音合成系统Debatts,用于生成反驳语音。Debatts接收两个语音提示:一方来自对手(提供辩论风格韵律),另一方来自说话人自身(保留身份特征)。系统在真实世界数据集上预训练,并引入额外的参考编码器来捕捉辩论风格。同时,研究者构建了一个辩论数据集以支持Debatts开发。实验表明,该系统在生成辩论风格语音方面优于传统零样本TTS方法。
原文摘要 · Abstract (English)
In debating, rebuttal is one of the most critical stages, where a speaker addresses the arguments presented by the opposing side. During this process, the speaker synthesizes their own persuasive articulation given the context from the opposing side. This work proposes a novel zero-shot text-to-speech synthesis system for rebuttal, namely Debatts. Debatts takes two speech prompts, one from the opposing side (i.e. opponent) and one from the speaker. The prompt from the opponent is supposed to provide debating style prosody, and the prompt from the speaker provides identity information. In particular, we pretrain the Debatts system from in-the-wild dataset, and integrate an additional reference encoder to take debating prompt for style. In addition, we also create a debating dataset to develop Debatts. In this setting, Debatts can generate a debating-style speech in rebuttal for any voices. Experimental results confirm the effectiveness of the proposed system in comparison with the classic zero-shot TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。