提出DINA框架,同时防御内部标签污染和外部攻击,提升NLP模型鲁棒性。
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
- 融合视觉领域噪声标签学习与对抗训练,统一防御双重威胁。
- 在真实游戏平台数据集上,模型准确率显著优于基线。
- 适合需高可靠性的客服与内容审核场景,推动AI公平部署。
随着大语言模型和生成式AI在客服与内容审核中的广泛应用,外部攻击和内部标签污染带来的对抗威胁日益突出。本文系统识别并解决这两类威胁,提出专为自然语言处理设计的统一防御框架DINA(Dual Defense Against Internal Noise and Adversarial Attacks)。该方法借鉴计算机视觉中的先进噪声标签学习技术,并结合对抗训练,同时缓解内部标签篡改与外部对抗扰动。在某在线游戏服务平台的真实数据集上进行的大量实验表明,相比基线模型,DINA显著提升了模型的鲁棒性与准确率。研究结果不仅强调了双重威胁防御的重要性,也为现实对抗场景下NLP系统的安全防护提供了可行策略,对实现公平、负责任的AI部署具有广泛意义。
原文摘要 · Abstract (English)
As large language models (LLMs) and generative AI become increasingly integrated into customer service and moderation applications, adversarial threats emerge from both external manipulations and internal label corruption. In this work, we identify and systematically address these dual adversarial threats by introducing DINA (Dual Defense Against Internal Noise and Adversarial Attacks), a novel unified framework tailored specifically for NLP. Our approach adapts advanced noisy-label learning methods from computer vision and integrates them with adversarial training to simultaneously mitigate internal label sabotage and external adversarial perturbations. Extensive experiments conducted on a real-world dataset from an online gaming service demonstrate that DINA significantly improves model robustness and accuracy compared to baseline models. Our findings not only highlight the critical necessity of dual-threat defenses but also offer practical strategies for safeguarding NLP systems in realistic adversarial scenarios, underscoring broader implications for fair and responsible AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。