自动生成韩语隐性攻击性语言数据,提升净化模型训练效果。
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
- 构建自动化管道生成含隐性攻击的韩语流行语
- 生成数据隐性攻击性更强,配对一致性高
- 可迁移至其他语言,适配简单微调的净化模型
语言净化旨在消除语言中的毒性。尽管中性-有毒配对数据集为训练净化模型提供了直接路径,但构建此类数据集面临两大挑战:一是需人工标注生成配对数据,二是攻击性词汇快速演变导致静态数据集迅速过时。为此,我们提出自动化配对数据生成管道K/DA,可生成具有隐性攻击性且符合潮流的韩语表达,使数据更适用于净化模型训练。实验表明,K/DA生成的数据在配对一致性上表现优异,隐性攻击性显著高于现有韩语数据集,且具备跨语言适用性。此外,该数据可支持通过简单指令微调训练出高性能净化模型。
原文摘要 · Abstract (English)
Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i) the need for human annotation to build paired data, and ii) the rapid evolution of offensive terms, rendering static datasets quickly outdated. To tackle these challenges, we introduce an automated paired data generation pipeline, called K/DA. This pipeline is designed to generate offensive language with implicit offensiveness and trend-aligned slang, making the resulting dataset suitable for detoxification model training. We demonstrate that the dataset generated by K/DA exhibits high pair consistency and greater implicit offensiveness compared to existing Korean datasets, and also demonstrates applicability to other languages. Furthermore, it enables effective training of a high-performing detoxification model with simple instruction fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。