首次系统研究语音离散表示中的口音信息,发现层选择影响最大。
Rethinking Discrete Speech Representation Tokens for Accent Generation
- 提出新评测框架,结合口音ABX任务与跨口音语音转换
- 深层表示保留口音信息更强,语音识别监督会削弱口音
- 减小码本大小无法有效分离口音与发音、说话人特征
离散语音表示(DSRTs)已成为语音生成的核心组件。尽管已有大量研究关注其语音和说话人信息,但口音信息在其中如何编码仍不清楚。本文首次系统性地探究了口音信息在DSRTs中的表征。我们提出了一个统一评估框架,通过新颖的口音ABX任务衡量口音信息的可访问性,并通过跨口音语音转换(VC)重建评估其可恢复性。基于该框架,我们分析了多种常用语音表示生成的DSRTs。结果表明:(1)层的选择对保留口音信息影响最大;(2)语音识别(ASR)监督显著降低口音信息;(3)简单缩小码本规模无法有效解耦口音与语音和说话人信息。
原文摘要 · Abstract (English)
Discrete Speech Representation Tokens (DSRTs) have become a foundational component in speech generation. While prior work has extensively studied phonetic and speaker information in DSRTs, how accent information is encoded in DSRTs remains largely unexplored. In this paper, we present the first systematic investigation of accent information in DSRTs. We propose a unified evaluation framework that measures both accessibility of accent information via a novel Accent ABX task and recoverability via cross-accent Voice Conversion (VC) resynthesis. Using this framework, we analyse DSRTs derived from several widely used speech representations. Our results reveal that: (1) choice of layers has the most significant impact on retaining accent information, (2) accent information is substantially reduced by ASR supervision; (3) naive codebook size reduction cannot effectively disentangle accent from phonetic and speaker information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。