用记忆机制提升长文本语音合成的连贯性与自然度
Long-Context Speech Synthesis with Context-Aware Memory

- 引入上下文感知记忆块,动态整合长期记忆与局部上下文
- 在段落级语音合成中实现更优的语调表现力与上下文一致性
- 适合需要高质量长文本语音生成的研究与应用
在长文本语音合成中,现有方法通常以句子为单位进行合成后拼接,忽略段落间上下文连贯性,导致语音自然度下降及风格、音色不一致。为此,我们提出基于上下文感知记忆(CAM)的长上下文文本转语音模型。CAM模块融合并检索长期记忆与局部上下文信息,支持长段落内动态记忆更新与传递,指导句级语音合成。此外,前缀掩码通过双向注意力机制增强前缀上下文学习能力,同时保持单向生成特性。实验表明,该方法在段落级语音合成中,在语调表现力、连贯性及上下文推理成本方面均优于基线与现有先进方法。
原文摘要 · Abstract (English)
In long-text speech synthesis, current approaches typically convert text to speech at the sentence-level and concatenate the results to form pseudo-paragraph-level speech. These methods overlook the contextual coherence of paragraphs, leading to reduced naturalness and inconsistencies in style and timbre across the long-form speech. To address these issues, we propose a Context-Aware Memory (CAM)-based long-context Text-to-Speech (TTS) model. The CAM block integrates and retrieves both long-term memory and local context details, enabling dynamic memory updates and transfers within long paragraphs to guide sentence-level speech synthesis. Furthermore, the prefix mask enhances the in-context learning ability by enabling bidirectional attention on prefix tokens while maintaining unidirectional generation. Experimental results demonstrate that the proposed method outperforms baseline and state-of-the-art long-context methods in terms of prosody expressiveness, coherence and context inference cost across paragraph-level speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。