针对韩语语音替换攻击,提出新型防御机制并提升检测鲁棒性。
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
- 利用韩语发音特性设计语音替换防御方法
- 通过混合语义-语音特征提升检测准确率
- 适合关注多语言对抗攻击的NLP研究者
随着恶意用户越来越多地使用语音替换规避仇恨言论检测,相关研究逐渐展开。然而仍存在两大挑战:一是现有研究忽视了韩语,而韩语因其音文字系统易受语音扰动影响;二是以往工作主要聚焦于数据集构建,缺乏架构级防御方案。为此,本文提出(1)基于韩语发音特性的拼音替换防御方法PHISH,以及(2)在模型架构层面融合语义与语音特征的MESH机制,以增强检测器鲁棒性。实验结果表明,所提方法在扰动和未扰动数据集上均表现优异,不仅提升了检测性能,还真实反映了恶意用户使用的对抗行为。
原文摘要 · Abstract (English)
As malicious users increasingly employ phonetic substitution to evade hate speech detection, researchers have investigated such strategies. However, two key challenges remain. First, existing studies have overlooked the Korean language, despite its vulnerability to phonetic perturbations due to its phonographic nature. Second, prior work has primarily focused on constructing datasets rather than developing architectural defenses. To address these challenges, we propose (1) PHonetic-Informed Substitution for Hangul (PHISH) that exploits the phonological characteristics of the Korean writing system, and (2) Mixed Encoding of Semantic-pHonetic features (MESH) that enhances the detector's robustness by incorporating phonetic information at the architectural level. Our experimental results demonstrate the effectiveness of our proposed methods on both perturbed and unperturbed datasets, suggesting that they not only improve detection performance but also reflect realistic adversarial behaviors employed by malicious users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。