通过细粒度引文关系追踪NLP、ML、CV领域研究演进路径。
How Does Research Evolve? Tracing Cross-Domain Trajectories in NLP, ML, and CV Through Claim-Grounded Typed Citations

- 基于论文中的具体论断标注引文类型,实现研究关系的精准溯源。
- 构建跨领域演进轨迹,发现视觉与大模型研究快速崛起,经典机器学习主题衰退。
- 验证了研究演进存在时间依赖性,内容与时间顺序共同决定知识发展脉络。
科研演进并非简单的知识累积。现有引用图常将各类引用简化为单一类型,难以刻画真实研究进展。本文提出SciTraj,一个面向自然语言处理、机器学习和计算机视觉领域的细粒度引用语料库,包含2015至2024年间32,559篇论文及573,126条带方向的六类研究关系边。每条边均关联其背后的论断句,并通过自然语言推理在局部上下文中验证其有效性。该语料库进一步将关系组织为多步类型的演化轨迹,追踪思想跨论文与时间的发展。评估显示:三名标注者一致性达Fleiss' κ=0.74,多数投票准确率79.9%,标签可靠。分析揭示各领域间研究流向存在明显壁垒;主题聚类识别出以视觉与大模型为核心的快速增长群组,以及若干传统机器学习主题的衰落趋势。在时间划分的链接预测基准和年份打乱的可证伪性测试中,SciTraj-Pair表现良好,但当发表年份随机打乱后AUC下降0.288,表明其预测不仅依赖内容,更依赖真实的时间演进顺序。
原文摘要 · Abstract (English)
How does research evolve, and can we trace it at the level of individual claims? Scientific progress is not simply a uniform accumulation of facts. Existing citation graphs usually collapse these roles into a single homogeneous edge type, limiting how we can analyze scientific progress. We introduce SciTraj, a typed citation corpus for tracing research evolution across natural language processing, machine learning, and computer vision. SciTraj includes 32,559 papers published between 2015 and 2024 and 573,126 directed edges spanning six research-relation types. Unlike traditional citation graphs, each edge is paired with the claim sentence that motivates its label. Claim-driven relations are verified by natural language inference against their local in-paper context. The corpus further organizes these relations into multi-step typed trajectories that trace how ideas develop across papers and over time. We evaluate the corpus along three dimensions. First, a three-annotator pilot achieves Fleiss' $κ=0.74$ and 79.9\% majority-vote precision for relation labels, indicating substantial agreement and reliable labeling. Second, corpus-level analyses reveal clear disciplinary siloing in the directional flow of research relations. Topic analysis further identifies rapidly growing clusters dominated by vision and LLM-related research and declining clusters associated with several classical machine-learning topics. We further evaluate SciTraj using a temporally split link-prediction benchmark and a year-shuffle falsifiability test that distinguishes genuine temporal signal from year-correlated content. Under this setting, \textsc{SciTraj-Pair} performs strongly, but its AUC drops by 0.288 when publication years are shuffled, showing that its predictions depend not only on content but also on the temporal order in which research develops.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。