首次研究语言模型对印地语等语言的自然扰动是否敏感,发现其抗性略强但仍有漏洞。
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
- 设计基于语言学规则的自然扰动攻击,模拟真实场景下的细微文本变化。
- 在多种印度语言上测试,发现语言模型对语言学扰动的敏感度低于非语言学攻击。
- 揭示语言模型在多语言、多语系环境中的脆弱性,适合关注低资源语言安全的研究者。
预训练语言模型(PLMs)已知对输入文本扰动敏感,但现有研究未聚焦于语言学基础的攻击,这类攻击更隐蔽且更贴近现实。本文首次系统研究语言模型是否对语言学基础扰动具有鲁棒性,涵盖多种印地语族语言及下游任务。结果表明,尽管语言模型对语言学扰动仍敏感,但相比非语言学攻击,其敏感度略有降低。这说明即使约束性强的攻击也具有效性。此外,研究覆盖不同语言家族和书写系统,揭示了跨语言影响的普遍性。
原文摘要 · Abstract (English)
Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, which are subtle and more prevalent in nature. In this paper, we study whether PLMs are agnostic to linguistically grounded attacks or not. To this end, we offer the first study addressing this, investigating different Indic languages and various downstream tasks. Our findings reveal that although PLMs are susceptible to linguistic perturbations, when compared to non-linguistic attacks, PLMs exhibit a slightly lower susceptibility to linguistic attacks. This highlights that even constrained attacks are effective. Moreover, we investigate the implications of these outcomes across a range of languages, encompassing diverse language families and different scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。