arXiv:2412.17947cs.CL2024-12中稿 · CHiPSAL Workshop a…被引 3

多语言印地语系文本中仇恨言论检测与目标识别,准确率达88.4%

IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages

  • 基于多语言Transformer的MultilingualRobertaClass模型
  • 子任务B准确率88.4%,子任务C准确率66.11%
  • 专为转写文本和多语言场景优化,适合社会安全研究者

本工作聚焦于印地语系语言(包括印地语、马拉地语、尼泊尔语、博杰普里语和梵语)中的仇恨言论检测与目标识别两个子任务。子任务B要求在在线文本中检测仇恨言论,子任务C则需识别仇恨言论的具体目标,如个人、组织或群体。我们提出MultilingualRobertaClass模型,基于预训练的多语言Transformer模型ia-multilingual-transliterated-roberta构建,针对多语言及转写文本的分类任务进行优化。该模型利用上下文嵌入处理语言多样性,并配备分类头实现二分类。在测试集上,子任务B获得88.40%的准确率,子任务C达到66.11%的准确率。

原文摘要 · Abstract (English)

This work focuses on two subtasks related to hate speech detection and target identification in Devanagari-scripted languages, specifically Hindi, Marathi, Nepali, Bhojpuri, and Sanskrit. Subtask B involves detecting hate speech in online text, while Subtask C requires identifying the specific targets of hate speech, such as individuals, organizations, or communities. We propose the MultilingualRobertaClass model, a deep neural network built on the pretrained multilingual transformer model ia-multilingual-transliterated-roberta, optimized for classification tasks in multilingual and transliterated contexts. The model leverages contextualized embeddings to handle linguistic diversity, with a classifier head for binary classification. We received 88.40% accuracy in Subtask B and 66.11% accuracy in Subtask C, in the test set.

仇恨言论检测多语言模型印地语系目标识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。