提出模态丢弃训练法,让多模态语音分离模型更抗模态缺失。
Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction
- 用模态丢弃训练提升多模态语音分离鲁棒性。
- 在LRS3数据集上,该方法在双说话人混合场景下性能稳定。
- 适合实际中目标说话人未提前注册的场景使用。
多模态目标说话人分离(MTSE)旨在利用音频和视觉等多源信息从语音混合中提取目标说话人。实际系统常因某模态主导而降低鲁棒性。本文研究训练策略与架构选择(尤其是归一化层)对鲁棒性的影。提出模态丢弃训练(MDT)优于标准训练和多任务训练(MTT)。在两说话人混合的LRS3数据集上,无论采用何种归一化层,MDT均表现稳定;而标准与MTT训练的模型易受模态主导影响,性能依赖于归一化方式。此外,经MDT训练的系统可直接使用提取语音作为注册信号,表明其适用于无预先注册的场景。
原文摘要 · Abstract (English)
The primary goal of multi-modal TSE (MTSE) is to extract a target speaker from a speech mixture using complementary information from different modalities, such as audio enrolment and visual feeds corresponding to the target speaker. MTSE systems are expected to perform well even when one of the modalities is unavailable. In practice, the systems often suffer from modality dominance, where one of the modalities outweighs the others, thereby limiting robustness. Our study investigates training strategies and the effect of architectural choices, particularly the normalization layers, in yielding a robust MTSE system in both non-causal and causal configurations. In particular, we propose the use of modality dropout training (MDT) as a superior strategy to standard and multi-task training (MTT) strategies. Experiments conducted on two-speaker mixtures from the LRS3 dataset show the MDT strategy to be effective irrespective of the employed normalization layer. In contrast, the models trained with the standard and MTT strategies are susceptible to modality dominance, and their performance depends on the chosen normalization layer. Additionally, we demonstrate that the system trained with MDT strategy is robust to using extracted speech as the enrollment signal, highlighting its potential applicability in scenarios where the target speaker is not enrolled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。