arXiv:2607.10191cs.SDcs.AI2026-07

用深度特征优化解决语音提取中质量与可懂度的矛盾

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

论文配图:Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization
图 1 · 摘自论文原文
  • 用WavLM特征作为优化锚点,防止因追求音质而牺牲可懂度
  • 在560毫秒流式处理下,词错误率从13.8%降至12.3%
  • 适合需要高可懂度的实时语音分离场景

面向目标说话人提取(TSE)的生成式流式模型普遍存在质量与可懂度之间的权衡问题:单纯优化感知音质会损害语音可懂度,反之亦然。我们发现该权衡并非源于流式架构限制,而是优化锚点选择不当所致。直接以音质指标为优化目标会引发灾难性奖励劫持,导致影响发音的关键内容被系统性删除以提升代理得分。为此,我们提出两项互补改进:扩大Conformer卷积核以增强局部时频建模能力,并采用基于WavLM的直接偏好优化(DPO)微调策略。通过WavLM余弦相似度对偏好对进行排序,该深度声学特征编码了音素结构与说话人身份信息,提供抗劫持的优化锚点。在560毫秒流式分块条件下,所提方法实现相对10.9%的可懂度提升(词错误率从0.138降至0.123),同时音质和说话人相似度略有提升。

原文摘要 · Abstract (English)

Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 to 0.123), with marginal simultaneous gains in audio quality and speaker similarity.

语音分离流式处理偏好优化可懂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。