arXiv:2409.19505cs.CL2024-09ACL被引 17

分析近2000篇NLP论文,揭示研究贡献的演变规律

The Nature of NLP: Analyzing Contributions in NLP Papers

  • 构建贡献分类体系并标注近2000篇论文摘要
  • 发现2010年后方法与数据集贡献显著上升,单篇论文贡献类型增多
  • 适合关注NLP发展脉络或文献综述的研究者

自然语言处理(NLP)是一个成熟且活跃的领域,但其研究本质仍存争议。本文通过量化分析NLP论文,提出一个研究贡献分类体系,并构建了包含近2000篇论文摘要的NLPContributions数据集,对科学贡献进行精细标注和分类。我们还提出一项新任务:自动识别论文中的贡献陈述并分类。实验结果显示,自1990年代以来,方法与数据集类贡献持续增长;而以语言和人类为中心的研究在1970-80年代占主导,1990-2000年代趋于下降,2010年代末重新回升。当前每篇论文平均贡献类型比以往更多。该数据集与分析为追踪研究趋势、生成数据驱动的文献综述提供了有力工具。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) is an established and dynamic field. Despite this, what constitutes NLP research remains debated. In this work, we address the question by quantitatively examining NLP research papers. We propose a taxonomy of research contributions and introduce NLPContributions, a dataset of nearly $2k$ NLP research paper abstracts, carefully annotated to identify scientific contributions and classify their types according to this taxonomy. We also introduce a novel task of automatically identifying contribution statements and classifying their types from research papers. We present experimental results for this task and apply our model to $\sim$$29k$ NLP research papers to analyze their contributions, aiding in the understanding of the nature of NLP research. We show that NLP research has taken a winding path -- with the focus on language and human-centric studies being prominent in the 1970s and 80s, tapering off in the 1990s and 2000s, and starting to rise again since the late 2010s. Alongside this revival, we observe a steady rise in dataset and methodological contributions since the 1990s, such that today, on average, individual NLP papers contribute in more ways than ever before. Our dataset and analyses offer a powerful lens for tracing research trends and offer potential for generating informed, data-driven literature surveys.

NLP研究贡献分析文献计量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。