人机协作生成藏文对抗文本,填补低资源语言安全研究空白
Human-in-the-Loop Generation of Adversarial Texts: A Case Study on Tibetan Script
- 通过人机协同设计,动态生成符合藏文特性的对抗文本
- 构建首个藏文对抗鲁棒性基准,验证模型在真实攻击下的脆弱性
- 为低资源语言安全研究提供可复用的方法框架,适合语言安全与评估研究者
基于深度神经网络的语言模型在各类自然语言处理任务中表现优异,但对文本对抗攻击仍极为脆弱。尽管对抗文本生成对NLP安全、可解释性、评估和数据增强至关重要,现有研究仍高度集中于英语,导致低资源语言缺乏高质量且可持续的对抗鲁棒性基准。主要挑战包括:1)因语言差异和资源匮乏,方法难以适配;2)自动化攻击易生成无效或歧义文本;3)模型持续演进,使旧有对抗样本失效。为此,我们提出HITL-GAT,一种基于通用策略的人机协同对抗文本生成系统。通过藏文案例研究,采用三种定制化生成方法,首次建立藏文对抗鲁棒性基准,为其他低资源语言提供重要参考。
原文摘要 · Abstract (English)
DNN-based language models excel across various NLP tasks but remain highly vulnerable to textual adversarial attacks. While adversarial text generation is crucial for NLP security, explainability, evaluation, and data augmentation, related work remains overwhelmingly English-centric, leaving the problem of constructing high-quality and sustainable adversarial robustness benchmarks for lower-resourced languages both difficult and understudied. First, method customization for lower-resourced languages is complicated due to linguistic differences and limited resources. Second, automated attacks are prone to generating invalid or ambiguous adversarial texts. Last but not least, language models continuously evolve and may be immune to parts of previously generated adversarial texts. To address these challenges, we introduce HITL-GAT, an interactive system based on a general approach to human-in-the-loop generation of adversarial texts. Additionally, we demonstrate the utility of HITL-GAT through a case study on Tibetan script, employing three customized adversarial text generation methods and establishing its first adversarial robustness benchmark, providing a valuable reference for other lower-resourced languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。