用强化学习提升语音伪造检测模型的泛化能力。
Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
- 采用群体相对策略优化(GRPO)进行强化学习微调。
- 在未见攻击数据上性能提升,且不降低原域表现。
- 负奖励机制可能是提升泛化的核心因素,适合安全研究者。
构建对未见攻击具有泛化能力的语音深度伪造检测模型仍是重大挑战。尽管领域已转向使用语音基础模型的预训练-微调范式,但多数方法仍仅依赖监督微调(SFT)。受大语言模型领域中强化学习微调的启发,本文探究了强化学习(特别是群体相对策略优化,GRPO)的影响。实验结果表明,纯基于GRPO的微调在跨域测试集上表现更优,同时保持目标域性能。该方法优于仅使用SFT或混合方案。消融研究表明,GRPO中的负奖励可能是性能提升的关键因素。
原文摘要 · Abstract (English)
Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuning paradigm using speech foundation models, most approaches rely solely on supervised fine-tuning (SFT). Inspired by the field of large language models, wherein reinforcement learning (RL) is used for model fine-tuning, we investigate the impact of RL, specifically Group Relative Policy Optimization (GRPO). The results from experiments using multiple detectors and test sets indicate that pure GRPO-based fine-tuning improves performance on out-of-domain test sets while maintaining performance on target-domain test data. This approach outperforms both SFT-only and hybrid setups. Our ablation studies further suggest that the negative reward in GRPO may be a key factor in this improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。