用提示词让大模型跨语言识别仇恨言论,效果比微调模型更泛化。
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
- 用零样本和少样本提示词测试大模型跨语言检测能力
- 在真实数据集上性能不如微调模型,但泛化能力更强
- 不同语言需定制提示词设计,效果差异明显
尽管自动化仇恨言论检测受到广泛关注,现有方法大多忽视了网络内容的语言多样性。多语言指令微调的大语言模型(如 LLaMA、Aya、Qwen、BloomZ)虽具备跨语言潜力,但其通过零样本和少样本提示词进行仇恨言论识别的有效性尚未充分探索。本研究评估了提示词驱动的检测方法在八种非英语语言上的表现,采用多种提示技术并与微调编码器模型对比。结果表明,尽管零样本和少样本提示在多数真实场景数据集上性能不及微调模型,但在功能测试中展现出更强的泛化能力。研究还发现提示词设计至关重要,每种语言往往需要定制化提示策略以实现最佳效果。
原文摘要 · Abstract (English)
Despite growing interest in automated hate speech detection, most existing approaches overlook the linguistic diversity of online content. Multilingual instruction-tuned large language models such as LLaMA, Aya, Qwen, and BloomZ offer promising capabilities across languages, but their effectiveness in identifying hate speech through zero-shot and few-shot prompting remains underexplored. This work evaluates LLM prompting-based detection across eight non-English languages, utilizing several prompting techniques and comparing them to fine-tuned encoder models. We show that while zero-shot and few-shot prompting lag behind fine-tuned encoder models on most of the real-world evaluation sets, they achieve better generalization on functional tests for hate speech detection. Our study also reveals that prompt design plays a critical role, with each language often requiring customized prompting techniques to maximize performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。