arXiv:2603.16070cs.CLcs.AI2026-03中稿 · ed被引 1

首个面向东南亚低资源语言的仇恨言论检测测试集

SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia

  • 构建跨四国语种的仇恨言论功能测试数据集,融合本地专家与大模型验证
  • 发现泰语、菲律宾语等模型准确率最低,尤其在俚语和隐晦表达上表现差
  • 揭示模型在反言辞识别上的不足,适合关注多语言内容安全的研究者

仇恨言论检测严重依赖语言资源,而这些资源主要集中在英语、中文等高资源语言,导致东南亚低资源语言的研究与平台工具开发面临障碍。为解决此问题,我们提出SEAHateCheck,首个专为印尼、泰国、菲律宾、越南设计的功能性测试集,涵盖印尼语、他加禄语、泰语和越南语。基于HateCheck框架并优化SGHateCheck方法,数据集通过大语言模型增强并经本地专家验证,确保文化相关性。实验显示,当前先进多语言模型在特定低资源语言中存在显著局限:他加禄语测试案例准确率最低,可能因语言复杂性和训练数据稀缺;以俚语为主的测试最难通过,因模型难以理解文化语境表达。诊断分析还暴露了模型在隐含仇恨识别和反言辞表达理解上的缺陷。作为首个针对这些东南亚语言的功能测试套件,SEAHateCheck为研究者提供了可靠基准,推动更具文化适配性的在线内容安全工具发展。

原文摘要 · Abstract (English)

Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low-resource languages in Southeast Asia, where diverse socio-linguistic contexts complicate online hate moderation. To address this, we introduce SEAHateCheck, a pioneering dataset tailored to Indonesia, Thailand, the Philippines, and Vietnam, covering Indonesian, Tagalog, Thai, and Vietnamese. Building on HateCheck's functional testing framework and refining SGHateCheck's methods, SEAHateCheck provides culturally relevant test cases, augmented by large language models and validated by local experts for accuracy. Experiments with state-of-the-art and multilingual models revealed limitations in detecting hate speech in specific low-resource languages. In particular, Tagalog test cases showed the lowest model accuracy, likely due to linguistic complexity and limited training data. In contrast, slang-based functional tests proved the hardest, as models struggled with culturally nuanced expressions. The diagnostic insights of SEAHateCheck further exposed model weaknesses in implicit hate detection and models' struggles with counter-speech expression. As the first functional test suite for these Southeast Asian languages, this work equips researchers with a robust benchmark, advancing the development of practical, culturally attuned hate speech detection tools for inclusive online content moderation.

仇恨言论低资源语言内容安全多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。