让语音模型删除特定数据,不重训也能保性能。
Speech Unlearning
- 提出语音数据删减方法,分样本和类别两种场景。
- 实验证明语音删减比图像文本更难,性能下降更明显。
- 适合关注隐私保护与模型可撤销性的研究者。
我们引入语音任务中的机器删减技术,这是一个新颖且未被充分探索的研究方向,旨在无需全量重新训练即可高效有效地移除特定数据对已训练语音模型的影响。该技术在隐私保护、清除过时或噪声数据以及缓解偏差方面具有重要意义。尽管机器删减已在计算机视觉和自然语言处理中有所研究,但由于语音数据具有高维、序列性和说话人依赖性,其在语音领域的应用仍几乎空白。我们定义了两个基础语音删减任务:样本删减(移除单个数据点,如一段语音录音)和类别删减(移除某一类数据,如某说话人的全部数据),同时保持剩余数据的性能。关键词检测和说话人识别实验表明,语音数据删减比图像或文本删减更具挑战性。最后,我们提出了若干未来方向,包括结构化训练、鲁棒评估、特征级删减、更广泛应用、可扩展方法及对抗鲁棒性。
原文摘要 · Abstract (English)
We introduce machine unlearning for speech tasks, a novel and underexplored research problem that aims to efficiently and effectively remove the influence of specific data from trained speech models without full retraining. This has important applications in privacy preservation, removal of outdated or noisy data, and bias mitigation. While machine unlearning has been studied in computer vision and natural language processing, its application to speech is largely unexplored due to the high-dimensional, sequential, and speaker-dependent nature of speech data. We define two fundamental speech unlearning tasks: sample unlearning, which removes individual data points (e.g., a voice recording), and class unlearning, which removes an entire category (e.g., all data from a speaker), while preserving performance on the remaining data. Experiments on keyword spotting and speaker identification demonstrate that unlearning speech data is significantly more challenging than unlearning image or text data. We conclude with key future directions in this area, including structured training, robust evaluation, feature-level unlearning, broader applications, scalable methods, and adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。