用代理监督训练语音分离模型,提升真实对话场景下的说话人提取效果。
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

- 构建7万+样本的多语言对话数据集,含语音混合、声纹注册和语音活动标签。
- 联合优化语音识别、说话人相似度、语音活动检测和音频感知质量四项指标。
- 仅微调分离器模块,在真实挑战赛中排名第二,说话人相似度与时间准确率最优。
针对真实对话混合语音中目标说话人提取(TSE)模型训练困难的问题,因缺乏大规模标注语料和纯净目标语音监督,本文提出PS4框架。首先,基于四个公开数据集构建包含71,771个样本的大规模训练语料库,涵盖中英文场景,每条样本包含重叠语音混合、各说话人注册音频、真实转录文本及帧级语音活动标签。其次,提出一种代理监督联合训练策略,通过四种可微分目标对基于BSRNN的TSE模型进行微调:语音识别交叉熵、说话人相似度、帧级语音活动检测和感知音频质量。仅更新分离器模块,从预训练检查点开始。在REAL-T挑战赛榜单上,PS4位列第二,各项指标中说话人相似度和时间F1表现最佳。
原文摘要 · Abstract (English)
Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。