arXiv:2501.14477eess.AScs.SD2025-01被引 6

用预训练模型联合优化语音清晰度与质量,提升目标说话人提取效果。

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

  • 基于Whisper模型构建生成式语音提取框架,融合语义与声学建模
  • 在多个基准上同时优于判别式与生成式基线,显著提升语音可懂度
  • 适合语音增强、会议记录等需高可懂度的场景

目标语音提取(TSE)旨在从多人重叠语音中分离出特定说话人的语音。现有方法多采用判别式策略,通常预测目标语音的时间-频率谱掩码,但掩码不完美常导致目标或非目标语音过抑制或欠抑制,影响听感质量。生成式方法通过混合语音和目标说话人线索重新合成目标语音,可实现更优的听感质量,但常忽略语音可懂度,导致语义内容被修改或丢失。受Whisper模型在目标说话人语音识别(ASR)中的成功启发,我们提出一种基于预训练Whisper模型的生成式TSE框架,将语义建模与基于流的声学建模联合优化,兼顾高可懂度与高感知质量。多基准实验表明,该方法在性能上全面超越现有判别式与生成式基线。演示音频见:https://aisaka0v0.github.io/GenerativeTSE_demo/

原文摘要 · Abstract (English)

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spectrogram mask for the target speech. However, imperfections in these masks often result in over-/under-suppression of target/non-target speech, degrading perceptual quality. Generative methods, by contrast, re-synthesize target speech based on the mixture and target speaker cues, achieving superior perceptual quality. Nevertheless, these methods often overlook speech intelligibility, leading to alterations or loss of semantic content in the re-synthesized speech. Inspired by the Whisper model's success in target speaker ASR, we propose a generative TSE framework based on the pre-trained Whisper model to address the above issues. This framework integrates semantic modeling with flow-based acoustic modeling to achieve both high intelligibility and perceptual quality. Results from multiple benchmarks demonstrate that the proposed method outperforms existing generative and discriminative baselines. We present speech samples on https://aisaka0v0.github.io/GenerativeTSE_demo/.

语音提取生成模型可懂度Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。