arXiv:2604.01562cs.SDcs.AI2026-04

研究普通话口音与克隆语音的差异,发现口音影响感知相似度和可懂度。

Acoustic and perceptual differences between standard and accented speech and their voice clones

  • 对比标准与重口音普通话及其克隆语音,结合计算与感知实验
  • 重口音说话者克隆语音的可懂度提升更明显,感知相似度更低
  • 建议将口音保留作为语音克隆中身份保真的关键指标

语音克隆常以整体质量评估,但对口音保留及其感知影响了解较少。本研究采用计算与感知相结合的方法,比较标准与重口音普通话及其克隆语音。基于嵌入的分析显示,在多个说话人判别嵌入空间中,重口音说话者的原声与克隆间距离更大,但经说话人内原声基线变异归一化后,该差异消失。感知实验表明,克隆语音对标准说话者的相似度评分更高,且从原声到克隆语音的可懂度均有提升,重口音语音的增益更显著。结果表明,口音差异会影响语音克隆中的感知身份匹配与可懂度,即使在基线归一化的嵌入距离中未体现,因此应将口音保留明确作为说话人身份保留的组成部分,而非依赖现成的说话人判别嵌入。

原文摘要 · Abstract (English)

Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences. We compare standard and heavily accented Mandarin speech and their voice clones using a combined computational and perceptual design. Embedding-based analyses showed larger original-clone distances for accented speakers in several speaker-discriminative embedding spaces, but this difference disappeared after normalizing against each speaker's within-original baseline variability. In the perception study, clones are rated as more similar to their originals for standard than for accented speakers, and intelligibility increases from original to clone, with a larger gain for accented speech. These results show that accent variation can shape perceived identity match and intelligibility in voice cloning even when it is not reflected in baseline-normalized speaker-embedding distance, and they motivate treating accent preservation as an explicit component of speaker identity preservation, rather than assuming that it is fully captured by off-the-shelf speaker-discriminative embeddings.

语音克隆口音保留感知实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。