arXiv:2605.20786cs.CL2026-05中稿 · ACL

阿拉伯语NLP二十年经验揭示:社会与制度问题比语言技术更难突破。

Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems

  • 从语言资源建设到社会计算,强调社区协作比任务本身更重要
  • 发现方言迁移失败、抑郁检测数据未进临床等三类关键教训
  • 适合关注低资源语言、社会计算与跨文化AI的研究者阅读

本文回顾了过去二十年阿拉伯语自然语言处理资源与研究基础设施的建设历程。尽管阿拉伯语有数亿使用者,却长期在资源投入上落后于英语或汉语。第一阶段聚焦基础语言学设施建设,第二阶段转向计算社会科学、社交媒体分析及社会导向应用。论文不重在罗列成果,而在于反思建设过程带来的启示。三个反直觉结论浮现:数据集构建本质上是社会过程;围绕共享任务形成的社区往往比任务本身更具价值;从语言资源向计算社会科学过渡时,暴露了传统NLP训练无法应对的挑战。文中指出三项失败:抑郁检测语料库未能进入临床实践、曾过度分散于过多共享任务而缺乏深度、长期误判现代标准阿拉伯语资源可直接迁移到方言任务。这些经历表明,服务弱势语言群体的最难关并非语言本身,而是社会、制度与认知层面的问题,亟需领域内少有教授的能力。

原文摘要 · Abstract (English)

This paper reflects on twenty years of building NLP resources and research infrastructure for Arabic, a language spoken by hundreds of millions yet historically underserved relative to languages such as English or Chinese. The first decade focused on foundational linguistic infrastructure; the second shifted toward computational social science, social media analysis, and socially oriented applications. Rather than cataloguing outputs, the paper examines what the experience of building them revealed. Three counterintuitive lessons emerge: building datasets is as much a social process as a technical one; communities formed around shared tasks often matter more than the tasks themselves; and moving from language resources to computational social science exposes challenges that traditional NLP training does not address. We discuss three failures: a depression detection corpus that never reached clinical practice, a period of spreading across too many shared tasks without sufficient depth, and a long-standing assumption that Modern Standard Arabic infrastructure would transfer cleanly to dialectal tasks. These experiences suggest that the hardest problems in developing NLP for underserved communities are not linguistic but social, institutional, and epistemic, and require competencies the field rarely teaches.

阿拉伯语NLP社会计算低资源语言语言公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。