arXiv:2606.26015cs.CL2026-06

为低资源语言塔塔尔语构建了首个文本净化系统,效果优于现有模型。

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

论文配图:The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
图 1 · 摘自论文原文
  • 基于塔塔尔语数据训练专用净化模型,提升本地化效果。
  • 新构建的塔塔尔语净化数据集支持后续研究与评估。
  • 跨语言迁移效果差,本土数据对性能至关重要。

文本净化是自动识别并缓解网络有害内容的关键技术,对保障在线社区安全至关重要。然而,低资源语言如塔塔尔语长期缺乏研究关注。本文提出 Tatoxa 系统,为塔塔尔语提供首个先进的文本净化方案。对比实验表明,该方法在关键质量指标上优于现有开源及商业大模型。同时,我们构建了一个专用于塔塔尔语文本净化的新型数据集,适用于低资源场景下的微调与评估。跨语言迁移实验显示,即使使用大规模俄语语料,从其他语言(包括文化相近的俄语)迁移的效果仍显著劣于在本族语塔塔尔语数据上直接训练的结果。

原文摘要 · Abstract (English)

Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention. In this paper we present Tatoxa, a novel state-of-the-art system for text detoxification in the Tatar language. Comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics. We also introduce a new dataset for text detoxification in Tatar, designed for fine tuning and evaluation in low resource settings. Finally, cross lingual transfer experiments indicate that transfer from other languages, including the culturally close Russian, performs significantly worse than training on native Tatar data even when a large Russian corpus is available.

文本净化低资源语言塔塔尔语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。