低词错误率误导了语音模型,纯语义标记无法生成清晰语音。
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

- 提出动态压缩分词器,精准对齐语义边界实现超低帧率
- 用纯语义标记生成语音时出现严重发音模糊,完全听不清
- 揭示语义分类与语音合成所需连续音素轨迹本质不兼容
追求语音理解与生成统一离散标记的语音语言模型社区,长期依赖词错误率(WER)作为表征质量的核心指标,尤其是基于Whisper风格分词器。这导致人们误认为低WER标记天然保留了可听语音合成所需信息。本文指出这是根本性误导:高频标记虽在生成任务中表现良好,但因隐含信息泄露,而极低帧率下的纯语义信息剥离了精细发音特征和微动态,破坏了基于常微分方程(ODE)的语音生成所需连续性。为验证此现象,需在不牺牲WER的前提下实现极端压缩,但传统固定步长下采样会随意截断音素边界。为此,我们设计了一种动态压缩分词器,智能对齐表示与语义边界,实现了超低帧率下的极低WER。利用这些分离出的“纯”语义标记,实验表明:即使采用理想持续时间对齐,生成语音仍出现严重发音模糊,听觉上不可识别。结果证明,低WER所奖励的语义分类与语音合成所需的连续音素轨迹本质上不兼容,彻底打破统一标记的幻象,主张显式解耦语音表示。
原文摘要 · Abstract (English)
The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) -- the core metric for Whisper-style tokenizers -- as the definitive proxy for representation quality. This fosters the assumption that low-WER tokens inherently preserve the information necessary for intelligible acoustic synthesis. We argue this is fundamentally deceptive. While high-frequency tokens succeed in generation tasks due to implicit information leakage, isolating pure semantic information at ultra-low frame rates strips away the finegrained articulation and micro-dynamics essential for ODE-based generation. Empirically validating this requires extreme compression without sacrificing WER -- a methodological bottleneck, as standard fixed-stride downsampling arbitrarily truncates phonetic boundaries. To overcome this, we develop a dynamic compression tokenizer that intelligently aligns representations with semantic boundaries, achieving ultra-low frame rates with exceptionally low WER. Using these isolated "pure" semantic tokens, we expose the WER trap: when conditioning generative models -- even with oracle duration alignments -- the reconstructed speech suffers from severe articulation blur and is rendered acoustically unintelligible. Our findings demonstrate that semantic categorization rewarded by low WER is inherently orthogonal to the continuous phonetic trajectories required for synthesis, shattering the illusion of the unified token and advocating for explicitly decoupled speech representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。