arXiv:2410.04690eess.AScs.LG2024-10被引 1

无需额外预测器,用分段隐式表示实现高效语音合成对齐

SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech

  • 用条件隐式神经表示将文本直接转为帧级特征
  • 自动定义分段边界,降低计算开销,零样本迁移表现优
  • 适合追求高效语音合成的开发者和研究者

我们提出SegINR,一种新型神经文语转换(TTS)方法,无需依赖辅助持续时间预测器或复杂的自回归(AR)/非自回归(NAR)帧级序列建模,即可实现序列对齐。SegINR通过优化的文本编码器提取嵌入,并利用条件隐式神经表示(INR)将每个文本嵌入转换为一段帧级特征,从而建模各段内的时序动态并自主确定分段边界,显著降低计算成本。该方法被集成到两阶段TTS框架中,用于语义标记预测。在零样本自适应TTS场景下的实验表明,SegINR在保持高语音质量的同时具备更高的计算效率,优于传统方法。

原文摘要 · Abstract (English)

We present SegINR, a novel approach to neural Text-to-Speech (TTS) that addresses sequence alignment without relying on an auxiliary duration predictor and complex autoregressive (AR) or non-autoregressive (NAR) frame-level sequence modeling. SegINR simplifies the process by converting text sequences directly into frame-level features. It leverages an optimal text encoder to extract embeddings, transforming each into a segment of frame-level features using a conditional implicit neural representation (INR). This method, named segment-wise INR (SegINR), models temporal dynamics within each segment and autonomously defines segment boundaries, reducing computational costs. We integrate SegINR into a two-stage TTS framework, using it for semantic token prediction. Our experiments in zero-shot adaptive TTS scenarios demonstrate that SegINR outperforms conventional methods in speech quality with computational efficiency.

语音合成隐式表示序列对齐TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。