arXiv:2412.17131cs.CL2024-12被引 4

用高效微调技术提升印地语和尼泊尔语仇恨言论检测能力

LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs

  • 采用参数高效微调(PEFT)降低大模型训练成本
  • 在印地语和尼泊尔语数据集上实现高精度识别
  • 为资源稀缺的天城文语言提供实用解决方案

仇恨言论检测在应对网络敌意及其现实后果方面日益重要。尽管近年来取得进展,针对使用天城文字母的语言的仇恨言论检测研究仍有限,因相关资源与工具稀缺。大型语言模型(LLMs)虽在语言任务中表现优异,但传统微调方式因模型规模过大而难以实施。本文提出基于参数高效微调(PEFT)的仇恨言论检测与目标识别方案。我们在Thapa等(2025)提供的天城文数据集上评估多个LLM,该数据集包含印地语和尼泊尔语两种语言的标注样本。实验结果表明,该方法在处理天城文内容方面具有显著有效性。

原文摘要 · Abstract (English)

The detection of hate speech has become increasingly important in combating online hostility and its real-world consequences. Despite recent advancements, there is limited research addressing hate speech detection in Devanagari-scripted languages, where resources and tools are scarce. While large language models (LLMs) have shown promise in language-related tasks, traditional fine-tuning approaches are often infeasible given the size of the models. In this paper, we propose a Parameter Efficient Fine tuning (PEFT) based solution for hate speech detection and target identification. We evaluate multiple LLMs on the Devanagari dataset provided by (Thapa et al., 2025), which contains annotated instances in 2 languages - Hindi and Nepali. The results demonstrate the efficacy of our approach in handling Devanagari-scripted content.

仇恨言论天城文高效微调多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。