新语音分词器同时保留语义与声学特征,提升多任务表现。
Speech Tokenizer is Key to Consistent Representation
- 融合语义与声学信息的新型分词方法
- 在语音编码等四类任务中显著提升表示质量
- 无需额外训练,适用于多种语音处理场景
语音分词在数字语音处理中至关重要,将连续语音信号转换为离散单元以支持各类计算任务。本文提出一种具有广泛适用性的新型语音分词器。尽管近期基于残差向量量化(RVQ)的方法已引入语义信息,但常忽略关键声学特征。我们提出一种先进方法,同步编码语言与声学信息,有效保留语调与情感内容。实证评估显示,该方法在语音编码、语音转换、情感识别及多模态语言建模等任务中均显著提升表示保真度,且无需额外训练。其通用性凸显其作为推动人工智能语音处理的关键工具潜力。
原文摘要 · Abstract (English)
Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicability across downstream tasks. While recent advances in residual vector quantization (RVQ) have incorporated semantic elements, they often neglect critical acoustic features. We propose an advanced approach that simultaneously encodes both linguistic and acoustic information, preserving prosodic and emotional content. Our method significantly enhances speech representation fidelity across diverse applications. Empirical evaluations demonstrate its effectiveness in speech coding, voice conversion, emotion recognition, and multimodal language modeling, without requiring additional training. This versatility underscores its potential as a key tool for advancing AI-driven speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。