首个面向韩语的去混淆与净化数据集,解决伪装毒语检测难题。
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
- 基于真实语料归纳韩语毒语混淆模式,构建可复用的变换规则
- 首次提供韩语毒语与去混淆后中性语对,支持双任务训练
- 适用于需处理韩语网络毒性内容的模型开发者与研究者
随着语言模型在在线环境中的广泛应用,毒性检测与净化日益重要。现有研究多聚焦非混淆文本,难以应对用户故意伪装的毒性表达。韩语因黏着性构词和韩文字形变异,极易被混淆。为此,我们提出KOTOX:首个面向韩语去混淆与净化的毒语数据集。将韩语混淆模式归类为语言学基础类别,基于真实案例定义变换规则,并开源配套转换工具包。利用这些规则,构建了包含毒语及其去混淆版本的成对语句。基于该数据集训练的模型在保持非混淆文本性能的同时,显著提升对混淆毒语的识别能力。这是首个同时支持去混淆与净化的韩语数据集,有助于提升韩语大模型对伪装毒语的防御能力。代码与数据已开源。
原文摘要 · Abstract (English)
As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention. Existing studies primarily focus on non-obfuscated text, which limits robustness when users intentionally disguise toxic expressions. In particular, Korean toxic expressions can be easily disguised through agglutinative morphology and Hangeul-specific orthographic variation. However, obfuscation in Korean remains largely unexplored, which motivates us to introduce a KOTOX: Korean toxic dataset for deobfuscation and detoxification. We categorize Korean obfuscation patterns into linguistically grounded classes, define transformation rules derived from real-world examples, and provide the resulting obfuscation framework as an open transformation package. Using these rules, we provide paired neutral and toxic sentences alongside their obfuscated counterparts. Models trained on our dataset better handle obfuscated text without sacrificing performance on non-obfuscated text. This is the first dataset that simultaneously supports deobfuscation and detoxification for the Korean language. We expect the dataset to facilitate better understanding and mitigation of obfuscated toxic content in LLM for Korean. Our code and data are available at https://github.com/leeyejin1231/KOTOX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。