研究离散语音符号对语调信息的编码能力,指导高效语音建模。
Benchmarking Prosody Encoding in Discrete Speech Tokens
- 通过人工修改语调测试离散符号的敏感性,评估其语调编码能力。
- 发现语调变化可被离散符号有效捕捉,但效果依赖于聚类数和预训练模型。
- 为语音语言模型设计离散表示提供实证指导,适合语音合成与理解研究者。
近年来,基于自监督学习(SSL)模型通过k-means聚类生成的离散符号,被广泛研究作为语音语言模型中的伪文本或各类任务的高效中间表示。然而,这些离散符号通常在语言模型或下游任务训练前独立学习,因此聚类方式(如使用的SSL模型、聚类数量)需凭经验选择。尤其对于语音语言模型,其不仅要理解语义,还需准确感知和生成包含语调特征的响应。但现有研究对离散符号编码语调的能力关注有限。为此,本研究通过系统分析离散符号对人工修改语调的敏感性,全面评估其语调编码性能,旨在为离散符号设计提供实用指导。
原文摘要 · Abstract (English)
Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks. However, these discrete tokens are typically learned in advance, separately from the training of language models or downstream tasks. As a result, choices related to discretization, such as the SSL model used or the number of clusters, must be made heuristically. In particular, speech language models are expected to understand and generate responses that reflect not only the semantic content but also prosodic features. Yet, there has been limited research on the ability of discrete tokens to capture prosodic information. To address this gap, this study conducts a comprehensive analysis focusing on prosodic encoding based on their sensitivity to the artificially modified prosody, aiming to provide practical guidelines for designing discrete tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。