arXiv:2412.10991cs.CLcs.AI2024-12被引 3

针对黎凡特阿拉伯语仇恨言论检测中的方言偏见与伦理难题,提出更包容的NLP方法。

Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection

  • 分析黎凡特阿拉伯语仇恨言论数据集的局限性与方言偏差
  • 指出公开多样数据稀缺导致模型泛化能力差
  • 呼吁构建文化敏感、上下文适配的仇恨言论检测工具

社交媒体已成为全球交流的核心,但也助长了仇恨言论的传播。对于黎凡特阿拉伯语等代表性不足的方言,仇恨言论检测面临独特的文化、伦理与语言挑战。本文探讨黎凡特阿拉伯语复杂的社会政治与语言环境,批判性审视当前仇恨言论检测所用数据集的局限性。研究揭示现有资源中公开、多样数据的匮乏,并分析方言偏见带来的后果。通过强调自然语言处理(NLP)工具需具备文化与上下文感知能力,主张在阿拉伯世界采用更细致、包容的仇恨言论检测策略。

原文摘要 · Abstract (English)

Social media platforms have become central to global communication, yet they also facilitate the spread of hate speech. For underrepresented dialects like Levantine Arabic, detecting hate speech presents unique cultural, ethical, and linguistic challenges. This paper explores the complex sociopolitical and linguistic landscape of Levantine Arabic and critically examines the limitations of current datasets used in hate speech detection. We highlight the scarcity of publicly available, diverse datasets and analyze the consequences of dialectal bias within existing resources. By emphasizing the need for culturally and contextually informed natural language processing (NLP) tools, we advocate for a more nuanced and inclusive approach to hate speech detection in the Arab world.

仇恨言论检测方言偏见阿拉伯语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。