构建保加利亚语毒性语言检测模型,兼顾敏感信息保护。
Detecting Toxic Language: Ontology and BERT-based Approaches for Bulgarian Text
- 构建保加利亚语毒性词本体,精准识别潜在有害内容。
- 基于4384条标注语句训练BERT模型,宏观F1达0.89。
- 适合需平衡内容过滤与敏感信息保护的平台使用。
在线交流中的毒性内容检测仍面临挑战,现有方案常误伤有价值信息,如医疗术语和少数群体相关文本。本文提出一种更精细的保加利亚语毒性内容识别方法,兼顾重要信息保留。研究采用两种不同方法:首先构建保加利亚语毒性词汇本体;其次建立包含4,384条手动标注句子的数据集,涵盖四类:毒性语言、医学术语、非毒性语言及少数群体相关词汇。在此基础上训练基于BERT的毒性分类模型,达到0.89的宏观F1分数。该模型可直接部署于真实环境,适合作为内容审核系统的核心组件。
原文摘要 · Abstract (English)
Toxic content detection in online communication remains a significant challenge, with current solutions often inadvertently blocking valuable information, including medical terms and text related to minority groups. This paper presents a more nu-anced approach to identifying toxicity in Bulgarian text while preserving access to essential information. The research explores two distinct methodologies for detecting toxic content. The developed methodologies have po-tential applications across diverse online platforms and content moderation systems. First, we propose an ontology that models the potentially toxic words in Bulgarian language. Then, we compose a dataset that comprises 4,384 manually anno-tated sentences from Bulgarian online forums across four categories: toxic language, medical terminology, non-toxic lan-guage, and terms related to minority communities. We then train a BERT-based model for toxic language classification, which reaches a 0.89 F1 macro score. The trained model is directly applicable in a real environment and can be integrated as a com-ponent of toxic content detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。