arXiv:2601.19781cs.SD2026-01中稿 · ICASSP 2026

提出一种兼顾语言与语调的语音离散化方法,适合语音大模型使用。

Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means

  • 用可微分K均值多任务微调语音标记,同时优化语音识别与重合成。
  • 在多个任务中保持语言和语调信息,同时去除说话人身份特征。
  • 适合对语调敏感的语音生成与理解任务,如语音大模型训练。

近年来,用离散标记表示语音成为热点,这些标记可作为语音语言模型(speechLMs)的伪文本或下游任务的高效中间表示。现有标记分为声学标记和音位标记:前者保留详细声学信息以用于重建,后者主要捕捉语言内容。然而在人类语音交流中,说话人信息等冗余声学细节被抽象,而语言和语调信息则被充分利用。因此,现有两类标记均不理想,尤其对语调敏感的任务(如speechLMs)。本文提出音位标记器(Phonological Tokenizer),通过可微分K均值对音位标记进行多任务微调,目标函数包含语音识别(ASR)与语音重合成。实验验证表明,该方法在多种任务中能有效保留语言与语调信息,同时适当消除说话人身份特征。

原文摘要 · Abstract (English)

In recent years, there has been growing interest in representing speech with discrete tokens, which serve as pseudo-text for speech language models (speechLMs) and as efficient intermediate representations for downstream tasks. These tokens are typically categorized as acoustic and phonetic tokens: the former holds detailed acoustic information for reconstruction while the latter mainly captures linguistic content. In human speech communication, however, unnecessary acoustic details such as speaker information are abstracted, while both linguistic and prosodic information are utilized for speech comprehension and production. Given this, neither type of token seems an ideal representation for tasks sensitive to prosody, such as speechLMs. In this study, we propose the Phonological Tokenizer, a method that fine-tunes phonetic tokens via differentiable k-means with a multi-task objective of ASR and speech resynthesis. Experimental validation on diverse tasks confirms that our tokens retain phonological (both linguistic and prosodic) information while appropriately discarding speaker identity.

语音建模离散化语调感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。