构建伊比利亚语仇恨言论检测多语言数据集与基准,填补低资源语言空白。
Bridging Gaps in Hate Speech Detection: Meta-Collections and Benchmarks for Low-Resource Iberian Languages
- 整合欧洲西班牙语数据集,统一标签与元数据,形成元数据集。
- 翻译并构建葡萄牙语、加利西亚语双变体对齐语料库,支持跨语言分析。
- 提供零样本、少样本及微调基准,适合低资源语言研究者使用。
仇恨言论对社会凝聚力与个体福祉构成严重威胁,尤其在社交媒体上迅速传播。尽管仇恨言论检测研究已取得进展,但主要集中在英语,导致低资源语言缺乏足够资源与基准。此外,许多语言存在多种方言变体,当前方法常忽略此因素。大型语言模型需大量数据才能可靠运行,而低资源语言往往难以满足。本文通过系统分析与整合现有资源,构建了欧洲西班牙语的仇恨言论数据集元集合,采用统一标签与元数据。进一步将该集合翻译为欧洲葡萄牙语,以及两种加利西亚语变体(分别趋近西班牙语和葡萄牙语),形成对齐的多语言语料库。基于这些资源,建立了伊比利亚语言仇恨言论检测的新基准。评估了先进大模型在零样本、少样本及微调设置下的表现,提供未来研究基线。同时开展跨语言分析。结果表明,多语言与方言感知方法对仇恨言论检测至关重要,为欧洲未充分代表语言的基准建设奠定基础。
原文摘要 · Abstract (English)
Hate speech poses a serious threat to social cohesion and individual well-being, particularly on social media, where it spreads rapidly. While research on hate speech detection has progressed, it remains largely focused on English, resulting in limited resources and benchmarks for low-resource languages. Moreover, many of these languages have multiple linguistic varieties, a factor often overlooked in current approaches. At the same time, large language models require substantial amounts of data to perform reliably, a requirement that low-resource languages often cannot meet. In this work, we address these gaps by compiling a meta-collection of hate speech datasets for European Spanish, standardised with unified labels and metadata. This collection is based on a systematic analysis and integration of existing resources, aiming to bridge the data gap and support more consistent and scalable hate speech detection. We extended this collection by translating it into European Portuguese and into a Galician standard that is more convergent with Spanish and another Galician variant that is more convergent with Portuguese, creating aligned multilingual corpora. Using these resources, we establish new benchmarks for hate speech detection in Iberian languages. We evaluate state-of-the-art large language models in zero-shot, few-shot, and fine-tuning settings, providing baseline results for future research. Moreover, we perform a cross-lingual analysis with our target languages. Our findings underscore the importance of multilingual and variety-aware approaches in hate speech detection and offer a foundation for improved benchmarking in underrepresented European languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。