基于训练动态设计多因素课程学习,提升目标说话人提取的实战性能。
Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction
- 联合调度信噪比、说话人数量等四类因素,渐进式训练复杂场景。
- 在多说话人场景下,相比随机采样提升显著,尤其在真实混合数据中。
- 通过可视化分析训练过程,自动识别易学/难学样本区域,指导课程设计。
目标说话人提取(TSE)旨在从多人混响语音中分离出特定说话人的声音。尽管基准测试表现优异,但实际应用中性能常因多种交互因素下降。现有课程学习方法通常单独处理这些因素,无法捕捉其复杂交互,且依赖预设难度指标,未必符合模型真实学习轨迹。为此,我们提出一种多因素课程学习策略,同步调度信噪比阈值、说话人数量、重叠比例及合成/真实数据比例,实现从简单到复杂的渐进学习。然而,如何在无预设假设下确定最优调度仍具挑战。因此,我们引入TSE-Datamap——一个基于训练动态的可视化框架,通过追踪各训练轮次中的置信度与变异性,揭示三类典型数据区域:(i) 易学样本,模型持续表现良好;(ii) 模糊样本,模型在不同预测间振荡;(iii) 难学样本,模型持续挣扎。基于此数据驱动洞察,所提方法优于随机采样,在复杂多说话人场景中取得显著提升。
原文摘要 · Abstract (English)
Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum learning approaches for TSE typically address these factors separately, failing to capture their complex interactions and relying on predefined difficulty factors that may not align with actual model learning behavior. To address this challenge, we first propose a multi-factor curriculum learning strategy that jointly schedules SNR thresholds, speaker counts, overlap ratios, and synthetic/real proportions, enabling progressive learning from simple to complex scenarios. However, determining optimal scheduling without predefined assumptions remains challenging. We therefore introduce TSE-Datamap, a visualization framework that grounds curriculum design in observed training dynamics by tracking confidence and variability across training epochs. Our analysis reveals three characteristic data regions: (i) easy-to-learn examples where models consistently perform well, (ii) ambiguous examples where models oscillate between alternative predictions, and (iii) hard-to-learn examples where models persistently struggle. Guided by these data-driven insights, our methods improve extraction results over random sampling, with particularly strong gains in challenging multi-speaker scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。