用大模型记忆强度衡量论文影响力,比引用数更及时准确。
LLM-Metrics: Measuring Research Impact Through Large Language Model Memory
- 通过大模型对论文的记性差异评估影响力,基于学术曝光度假设。
- 17个模型中9个显著相关,与引用数相关性达0.1495(p<0.0004)。
- 小模型反而更准,支持信息筛选机制,适合跨学科实时评估。
引用次数仍是评估研究影响力的主流指标,但存在时间滞后、学科偏差和马太效应等缺陷。本文提出LLM-Metrics,一种基于大语言模型(LLMs)参数记忆的研究影响力评估方法。核心假设是高影响力论文在学术界曝光更多,其文本被纳入训练数据,使模型形成更强的参数记忆。设计了四类多项选择探针(标题、作者、方法、会议识别),评估2023–2024年发表的549篇计算机科学论文,覆盖17个来自六家厂商的模型(参数量0.5B至72B)。17个模型中有15个呈现正向预测,其中9个在p<0.05水平显著,整体与引用数的斯皮尔曼相关系数为rho=0.1495(p=0.0004)。三项额外发现支持该机制:一是2024年论文相关性更强(rho=0.1880),当时引用数接近零,排除反向因果;二是作者识别探针区分力最强,符合暴露驱动记忆;三是模型规模与预测能力非单调,3B参数的Llama-3.2-3B-Instruct(rho=0.1829)优于多数更大模型,支持小模型作为有效信息过滤器的假设。LLM-Metrics提供了一种实时、跨学科、不依赖引用的评估范式。
原文摘要 · Abstract (English)
Citation counts remain the dominant metric for assessing research impact, yet they suffer from well-documented limitations: temporal lag, disciplinary bias, and Matthew effects. Here we propose LLM-Metrics, a research-impact assessment metric derived from the parametric memory of large language models (LLMs). The central hypothesis is that high-impact papers receive greater exposure in the academic community, that this exposure enters LLM training data in textual form, and that models consequently form stronger parametric memory of these papers. We designed four types of multiple-choice probes, covering title recognition, author recognition, method recognition, and venue recognition, and evaluated 549 computer science papers published in 2023-2024 across 17 LLMs spanning 0.5B to 72B parameters from six vendors. Of the 17 models, 15 produced positive predictions, 9 of which were significant at p less than 0.05, with an overall Spearman correlation of rho = 0.1495 and p = 0.0004 against citation counts. Three additional findings support the proposed mechanism. First, the predictive signal was stronger for 2024 papers, rho = 0.1880, whose citation counts were near zero at model-training time, reducing the plausibility of a simple reverse-causality explanation. Second, author-recognition probes showed the strongest discriminative power, consistent with an exposure-driven memory mechanism. Third, model scale and predictive power were non-monotonic: a 3B-parameter model, Llama-3.2-3B-Instruct, with rho = 0.1829, outperformed most larger models, supporting a selective-memory hypothesis in which the limited capacity of smaller models can serve as an effective information filter. LLM-Metrics offers a real-time, cross-disciplinary, citation-independent paradigm for research assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。