arXiv:2608.23421cs.CL2026-08

分析7120篇阿拉伯语NLP论文,揭示研究趋势与短板。

A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study

论文配图:A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
图 1 · 摘自论文原文
  • 用主题建模和引用分析挖掘研究热点与趋势
  • 2020年后论文激增,沙特、美国、埃及产出最多
  • 发现马格里布、伊拉克等方言研究严重不足

过去十年间,阿拉伯语自然语言处理(NLP)因阿拉伯世界数字化转型、社交媒体和大语言模型(LLMs)的推动迅速发展。然而,尚缺乏全面的量化元分析。本研究对1960至2026年间来自arXiv、ACL Anthology、Semantic Scholar、Crossref、OpenAlex及额外定向OpenAlex子集的7,120篇阿拉伯语NLP论文进行文献计量与主题分析。采用BERTopic进行主题建模,结合回归分析、社会网络分析与地理制图。结果表明,2020年后发表量显著上升,主要受变压器模型与大语言模型驱动。主题建模识别出19个主题,最大主题涵盖文本、语音、翻译与识别,共2,942篇。引文分析显示论文年龄与被引次数呈正相关(r = 0.245, p < 0.001),回归分析(R² = 0.105)表明在OpenAlex或Semantic Scholar索引及机构隶属关系与更高被引相关。沙特阿拉伯、美国和埃及为研究产出领先国家。任务-方言差距矩阵揭示了马格里布、伊拉克与苏丹方言的摘要生成等方向研究薄弱。最大主题具有最高H指数(90),其次为情感分析(57)。该定量方法补充现有定性综述,并建议优先支持资源匮乏方言并开发文化适配基准。

原文摘要 · Abstract (English)

Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis remains absent. This study presents a bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from five platforms (arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex) plus an additional targeted OpenAlex subset. We employ BERTopic for topic modeling, regression analysis, social network analysis, and geographic mapping. Our findings show a significant publication surge after 2020, driven by transformer models and LLMs. Topic modeling identifies 19 themes, the largest centered on text, speech, translation, and recognition (2,942 papers). Citation analysis reveals a positive correlation between paper age and citations (r = 0.245, p < 0.001); regression (R^2 = 0.105) shows that indexing in OpenAlex or Semantic Scholar and institutional affiliation are associated with higher citations. Saudi Arabia, the United States, and Egypt lead in research output. A task-dialect gap matrix identifies understudied areas, including summarization for Maghrebi, Iraqi, and Sudanese dialects. The largest topic has the highest H-index (90), followed by sentiment analysis (57). Our quantitative approach complements existing qualitative surveys and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks.

阿拉伯语NLP文献计量研究缺口方言研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。