让对话语音合成更自然,支持两人同时说话的重叠与换话。
KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction
- 用音素栅格控制每人的语音生成,分通道输出双人对话。
- 用户评测显示语音自然度更高,换话更频繁,重叠更真实。
- 适合需要逼真多人交互语音的应用,如虚拟助手、智能客服。
实现全双工语音对话需要大量双通道、单人一通道的会话语音数据。尽管已有会话语音合成(TTS)系统,但对回话信号、打断和重叠等双人同时发声现象缺乏鲁棒性。为此,本文提出KABURI-TTS:基于音素键控的双通道会话语音生成方法。该方法以每个说话人的音素栅格为输入,在帧级上结合音素与语音活动检测结果,分别生成两通道的语音。由于音素栅格由独立模块提供,该方法可实现可控的单人一通道双人对话合成。用户评估表明,相比强基线模型,所提方法在语句与交互层面均显著提升自然度。语音活动分析进一步验证了其生成更多重叠和更频繁换话的能力。
原文摘要 · Abstract (English)
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps that occur while the interlocutor is speaking. In this work, aiming at conversational speech synthesis that reproduces human-like overlap, we propose KABURI-TTS. KABURI-TTS takes a per-speaker phoneme raster as input and renders the speech of the two speakers on separate channels, conditioned on the per-frame phonemes and the voice activity derived from them. Because the phoneme raster is supplied by a separate module, the proposed method enables controllable generation of one-speaker-per-channel, two-party spoken dialogue. A user evaluation shows that, compared with strong baselines, the proposed method attains higher naturalness at both the utterance and the interaction level. Furthermore, an analysis of voice activity confirms that the proposed method produces more overlap and more frequent turn-taking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。