arXiv:2607.05401cs.DLcs.AI2026-07

分析3.6万篇ICLR论文,发现评审分数与未来影响力几乎无关。

Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025

论文配图:Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025
图 1 · 摘自论文原文
  • 用语义分析模型识别出推动研究方向的'催化剂'论文
  • 话题开创者使新领域引用量增长7.55倍,跨领域桥梁达11.52倍
  • 同行评审得分完全无法预测论文是否成为研究转折点

过去十年中,词向量、Transformer、大规模预训练和人类反馈强化学习等少数方法深刻改变了NLP与AI研究。如今ICLR对所有投稿公开评分与录用结果。但这些评审信号能否在提交时就识别出将改变研究轨迹的论文,尚未在大规模数据上验证。本文基于2017至2025年共36,113篇ICLR论文,识别出‘催化剂’:其后续研究显著转向的论文。比较四种颠覆性度量方法(CD指数、node2vec、EDM、LLM语义评分器),定义五类催化剂分类(话题开创者、话题桥梁、同主题重定向者、同步型、认知错配型)。EDM在识别高被引论文方面表现最佳(AUC 0.83),显著优于其他方法。话题开创者使新话题引用占比增长7.55倍,话题桥梁则带来11.52倍的跨领域引用流动。结果显示,同行评审分数与未来颠覆性几乎无关(|ρ|≤0.005),录用与拒稿论文的EDM均值无显著差异(p=0.11)。

原文摘要 · Abstract (English)

A small number of methodological contributions, including word2vec, the Transformer, large-scale pre-training, and reinforcement learning from human feedback, have reshaped NLP and AI research over the past decade. OpenReview now makes numeric reviewer scores and accept/reject decisions public for every ICLR submission. Whether such review signals identify trajectory-changing papers at submission time, however, remains untested at corpus scale. We answer this question on $36{,}113$ papers from ICLR 2017--2025, identifying \emph{catalysts}: papers whose descendants measurably redirect future research. We compare four disruptiveness measures (the Consolidation/Destabilization (CD) index, node2vec, the direction-aware Embedding Disruptiveness Measure (EDM), and an LLM-based semantic rater) and define a five-type operational catalyst taxonomy (topic initiator, topic bridge, within-topic redirector, simultaneous, and recognition-misaligned). EDM leads at identifying highly cited ICLR papers (AUC $0.83$ vs.\ $0.60$ for CD, $0.49$ for node2vec, and $0.42$ for the LLM rater). Topic initiators precede a $7.55{\times}$ topic-share growth and topic bridges precede an $11.52{\times}$ growth in cross-topic citation flow versus year-matched controls. We found that the peer review scores are essentially orthogonal to future disruptiveness ($|ρ|{\leq}0.005$; accepted and rejected papers have indistinguishable mean EDM, $p{=}0.11$).

论文评估AI趋势评审机制影响力预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。