arXiv:2603.26791cs.DLcs.AI2026-03

用大模型联合评估论文引用影响力,更准且省成本。

Crystal: Characterizing Relative Impact of Scholarly Publications

  • 用大模型联合排名一篇论文引用的所有文献,避免孤立分析。
  • 比现有最佳方法准确率高9.5%,F1高8.3%。
  • 适合需要批量分析引用影响力的科研评估者使用。

评估被引论文的影响力通常只分析其在引用论文中的孤立上下文。这忽略了论文引用的全部背景,难以进行相对比较。我们提出Crystal,通过大语言模型(LLMs)联合排序一篇论文所引用的所有文献。为缓解LLM的顺序偏差,每篇列表随机排列三次,通过多数投票聚合影响力标签。该联合方法利用完整引用上下文,更可靠地区分关键参考文献。Crystal在人工标注的引用数据集上,准确率提升9.5%,F1提升8.3%。同时减少LLM调用次数,效率更高;使用开源权重模型仍优于基线,支持可扩展、低成本的引用影响力分析。对ACL年度经典论文的案例研究显示,Crystal的影响力判断与长期科学认可高度一致。我们发布了Crystal-Bank数据集,包含46.8k篇论文的排名与影响力标签及代码。

原文摘要 · Abstract (English)

Assessing a cited paper's impact is typically done by analyzing its citation context in isolation within the citing paper. While this focuses on the most directly relevant text, it prevents relative comparisons across all the works a paper cites. We propose Crystal, which instead jointly ranks all cited papers within a citing paper using large language models (LLMs). To mitigate LLMs' positional bias, we rank each list three times in a randomized order and aggregate the impact labels through majority voting. This joint approach leverages the full citation context, rather than evaluating citations independently, to more reliably distinguish impactful references. Crystal outperforms a prior state-of-the-art impact classifier by +9.5% accuracy and +8.3% F1 on a dataset of human-annotated citations. Crystal further gains efficiency through fewer LLM calls and outperforms prior baselines using an open-weight model, enabling scalable, cost-effective citation impact analysis. In a case study of ACL Test-of-Time award-winning papers, we find that Crystal's impact characterizations align closely with long-term scientific recognition. We release Crystal-Bank, a 46.8k-paper dataset with rankings and impact labels, along with code.

引用分析大模型应用影响力评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。