arXiv:2509.15626cs.SDeess.AS2025-09中稿 · INTERSPEECH 2026

首个公开语音印象控制数据集,实现精准音色风格调节。

LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control

论文配图:LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
图 1 · 摘自论文原文
  • 用双语句分离说话人与音色印象特征
  • 控制误差降低至0.41(客观)和0.92(主观)
  • 适合需要精细音色调节的语音合成研究

数值化语音印象(VI)控制(如调节明亮度)可实现文本到语音(TTS)的细粒度调节。然而面临两大挑战:缺乏公开数据集及印象泄露问题——参考音频会使生成语音偏离目标印象。为此,我们提出首个基于LibriTTS-R的公开VI数据集LibriTTS-VI。针对泄露问题,我们假设单个参考音频会混淆说话人身份与印象特征。为此提出:1)使用同一说话人的两个语句进行解耦训练,分别用于说话人与印象条件;2)无参考方法,仅通过目标印象参数控制音色。实验表明,最优方法使11维VI均方误差从0.61降至0.41(客观),主观评估从1.15降至0.92。相较提示式TTS,本方法克服了控制不精确及印象与语义纠缠的问题。

原文摘要 · Abstract (English)

Numerical voice impression (VI) control (e.g., scaling brightness) enables fine-grained control in text-to-speech (TTS). However, it faces two challenges: no public corpus and impression leakage, where reference audio biases synthesized voice away from the target VI. To address the first challenge, we introduce LibriTTS-VI, the first public VI corpus built on LibriTTS-R. For the second, we hypothesize a single reference causes leakage by entangling speaker identity and VI. To mitigate this, we propose 1) disentangled training with two utterances from the same speaker for speaker and VI conditioning, and 2) a reference-free method controlling the impression solely via target VI. Experimentally, our best method improves controllability: 11-dimensional VI mean squared error drops from 0.61 to 0.41 objectively and 1.15 to 0.92 subjectively. A comparison with a prompt-based TTS reveals imprecise numerical control and entanglement between VI and text semantics, which our methods overcome.

语音合成印象控制数据集解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。