构建论文观点演化图谱,追踪NLP领域观点的延续与修正。
ClaimFlow: Tracing the Evolution of Scientific Claims in NLP
- 以观点为单位构建跨论文关联网络,明确支持、扩展、反驳等关系。
- 发现63.5%的观点从未被重用,仅11.1%被挑战,多数观点被修正而非推翻。
- 适用于研究观点演进、模型评估或文献综述的学者。
科学论文提出观点,后续工作可能支持、扩展或反驳这些观点。现有引文分析方法仅捕捉对话的片段。本文构建了基于1,617篇ACL论文(1979–2025)的《ClaimFlow》框架,人工标注5,689个观点及4,871条跨论文观点关系,涵盖支持、扩展、质疑、反驳和背景引用五类。在此基础上,提出新任务“观点关系分类”,要求模型从文本与引文上下文中推断对被引观点的科学立场。在该任务上,神经模型与大模型取得0.81宏F1基线,表明任务可解但仍有提升空间。进一步将框架扩展至约1.3万篇论文,分析几十年间观点演化:63.5%的观点未被重用,仅11.1%被挑战;广泛传播的观点更多通过修正与扩展演变,而非直接支持或驳倒。总体而言,ClaimFlow为观察NLP思想如何演进提供了新视角。
原文摘要 · Abstract (English)
Scientific papers advance $\textit{claims}$ that later work supports, extends, or sometimes refutes. Yet existing methods for citation and claim analysis capture only fragments of this dialogue. In this work, we make these interactions explicit at the level of individual scientific claims. We introduce $\texttt{ClaimFlow}$, a claim-centric view of the NLP literature, built from $1{,}617$ ACL Anthology papers $(1979 - 2025)$ that are manually annotated with $5{,}689$ claims and $4{,}871$ cross-paper claim relations, indicating whether a citing paper $\texttt{supports}$, $\texttt{extends}$, $\texttt{qualifies}$, $\texttt{refutes}$, or references a cited claim as $\texttt{background}$. Building on $\texttt{ClaimFlow}$, we define a new task -- $\textit{Claim Relation Classification}$ -- which requires models to infer the scientific stance toward a cited claim from the text and citation context. Evaluating neural models and large language models on this task, we report baseline performance of $0.81$ macro-F1, suggesting that the task is tractable while leaving room for improvement. We then scale this framework to $\sim$$13k$ NLP papers to study claim evolution across decades of NLP research. We show that $63.5\%$ claims are never reused; only $11.1\%$ are ever challenged. Widely propagated claims are more often $\textit{reshaped}$ through qualification and extension than supported or refuted. Overall, $\texttt{ClaimFlow}$ offers a lens for examining how ideas shift and mature within NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。