用自净化方法训练流匹配模型,自动过滤错误标签数据。
Training Flow Matching Models with Reliable Labels via Self-Purification
- 训练中利用模型自身识别可疑数据,无需额外模块
- 在含噪标签下仍能生成精准条件样本,性能优于基线
- 适用于真实场景语音数据,对噪声标签有强鲁棒性
训练数据普遍存在标签错误,源于人工标注误差、标签模型局限等噪声源,严重影响模型性能。本文提出自净化流匹配(SPFM),在流匹配框架内通过模型自身动态识别可疑数据,无需预训练模型或额外模块。实验表明,使用SPFM训练的模型即使在含噪标签下,仍能准确生成指定条件的样本。此外,在包含真实场景语音的TITW数据集上验证了SPFM的鲁棒性,性能超越现有基线。
原文摘要 · Abstract (English)
Training datasets are inherently imperfect, often containing mislabeled samples due to human annotation errors, limitations of tagging models, and other sources of noise. Such label contamination can significantly degrade the performance of a trained model. In this work, we introduce Self-Purifying Flow Matching (SPFM), a principled approach to filtering unreliable data within the flow-matching framework. SPFM identifies suspicious data using the model itself during the training process, bypassing the need for pretrained models or additional modules. Our experiments demonstrate that models trained with SPFM generate samples that accurately adhere to the specified conditioning, even when trained on noisy labels. Furthermore, we validate the robustness of SPFM on the TITW dataset, which consists of in-the-wild speech data, achieving performance that surpasses existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。