arXiv:2608.21408cs.AI2026-08

对比多种方法,提升低资源罗马乌尔都语仇恨言论识别效果

Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

  • 用LoRA等参数高效技术微调大模型,降低计算成本
  • 提示工程在少量样本下表现优于传统微调方法
  • 适合资源有限但需精准识别方言仇恨内容的团队

由于互联网和社交媒体的普及,有毒及仇恨内容呈指数级增长,造成显著心理伤害与社会负面影响。罗马乌尔都语作为巴基斯坦及全球乌尔都语社群使用的低资源语言,因语法非正式、句式不统一、拼写多样而面临更大挑战。本研究旨在探索在数据有限条件下,最有效的仇恨言论分类技术。通过四组实验对比:1)零样本直接推理,评估大模型对罗马乌尔都语的理解能力;2)使用LoRA进行参数高效微调(PEFT),仅更新少量参数以降低算力开销;3)采用混合与人工设计提示的提示调优,仅用极小训练集实现高效学习;4)基于精心设计指令提示的零样本与少样本提示工程,无需额外训练。结果表明,在低资源环境下,提示工程与参数高效微调结合方案表现最优。

原文摘要 · Abstract (English)

Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts. Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings. This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data. To address this, the study investigates and compares the latest approaches, in-cluding prompt tuning, parameter-efficient fine-tuning (PEFT) using LoRA, and prompt en-gineering, under various experimental configurations. To achieve this objective, four exper-iments were designed. The first experiment involved direct inferencing with LLMs without any fine-tuning, to evaluate how well these models understand Roman Urdu in a zero-shot setting, especially given limited data. The second experiment utilized parameter-efficient fine-tuning (PEFT) with LoRA, which updates only a small subset of parameters, thereby reducing computational cost. The third experiment explored prompt tuning with both mixed and manually crafted prompts, using very small sets of training examples relative to the entire dataset, making it computationally efficient as well. Finally, the fourth experiment applied prompt engineering through zero-shot and few-shot learning, relying solely on care-fully designed instruction prompts for classification without further training.

仇恨言论低资源语言提示工程LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。