系统梳理非洲NLP二十年研究进展与贡献者画像
The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysis
- 基于2.2千篇论文构建非洲NLP贡献数据集
- 分析4.9千作者与7.8千标注贡献句,揭示研究趋势
- 开源工具助力数据驱动的NLP研究探索
自然语言处理(NLP)正因大语言模型持续演进,研究与实践每日突破。本文通过分析2005至2025年共2.2千篇非洲NLP(AfricaNLP)论文、4.9千名贡献作者及7.8千条人工标注的贡献语句(AfricaNLPContributions),探讨非洲NLP在论文数量、研究主题、任务类型、数据、方法与任务创新等方面的进展。研究还聚焦作者、机构与资助方的贡献情况,提供基准结果与可视化探索工具。该数据集与研究探索工具已开源,可为追踪非洲NLP发展轨迹与推动数据驱动研究提供支持。
原文摘要 · Abstract (English)
Natural Language Processing (NLP) is undergoing constant transformation, as Large Language Models (LLMs) are driving daily breakthroughs in research and practice. In this regard, tracking the progress of NLP research and automatically analyzing the contributions of research papers provides key insights into the nature of the field and the researchers. This study explores the progress of African NLP (AfricaNLP) by asking (and answering) research questions about the progress of AfricaNLP (publications, NLP topics, and NLP tasks), contributions (data, method, and task), and contributors (authors, affiliated institutions, and funding bodies). We quantitatively examine two decades (2005 - 2025) of contributions to AfricaNLP research, using a dataset of 2.2K NLP papers, 4.9K contributing authors, and 7.8K human-annotated contribution sentences (AfricaNLPContributions), along with benchmark results. Our dataset and AfricaNLP research explorer tool will provide a powerful lens for tracing AfricaNLP research trends and holds potential for generating data-driven research approaches. The resource can be found in GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。