将朗读语音转为自然对话语音,提升虚拟助手真实感
Bridging the Gap: Converting Read Text to Conversational Dialogue
- 用深度网络分析并调整语调、重音、节奏等韵律特征
- 在多个数据集上提升语音自然度,MOS评分显著提高
- 适合语音助手、客服系统等需真实对话感的应用
近年来,将朗读语音转换为自然对话语音成为语音处理领域的重要研究方向。核心挑战在于保持语音自然性和可懂性的同时,降低实时应用中的计算开销。传统朗读语音缺乏自然对话所需的细微韵律变化,限制了其在虚拟助手、客户服务和语言学习工具中的应用。本文提出一种新方法——基于对话上下文的韵律调整(PACC),通过深度神经网络分析并修改语调、重音与节奏等韵律特征。不同于传统方法,本研究采用高保真生成对抗网络(HiFi-GAN)进行语音合成。实验表明,该方法在多个语音数据集上显著提升转换效果,增强语音自然度,并在模型准确率与平均意见分(MOS)评估中建立新基准。研究还证明该方法可拓展至其他语音转换场景。
原文摘要 · Abstract (English)
In recent advancements within speech processing, converting read speech to conversational speech has gained significant attention. The primary challenge in this domain is maintaining naturalness and intelligibility while minimizing computational overhead for real-time applications. Traditional read speech often lacks the nuanced prosodic variation essential for natural conversational interactions, posing challenges for applications in virtual assistants, customer service, and language learning tools. This paper introduces a novel approach, Prosodic Adjustment with Conversational Context (PACC), aimed at converting read speech into natural conversational speech used in various modern applications. PACC utilizes advanced deep neural networks to analyze and modify prosodic features such as intonation, stress, and rhythm. Unlike conventional methods, our approach uses High-Fidelity Generative Adversarial Networks (HiFi-GAN) for speech synthesis. Our experimental results demonstrate significant improvements in speech conversion, enhancing naturalness and achieving better model accuracy with additional training on speech datasets. This research establishes new benchmarks in speech conversion tasks and Mean Opinion Score (MOS) evaluation for testing model accuracy, and we show that our approach can be successfully extended to other speech conversion applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。