用领域感知检索和隐式推理提升论文评估准确率
Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning
- 引入领域感知检索,获取相关同期研究以判断创新性
- 通过隐式推理机制实现对方法与动机的深度理解
- 在真实推荐系统中验证效果,单篇最高获超万次浏览
随着学术论文数量激增,识别高质量研究日益困难。现有基于大语言模型(LLM)的论文评估方法常受限于过时的领域知识和有限的推理能力。本文提出PaperEval框架,通过两个关键组件克服这些问题:1)领域感知论文检索模块,获取相关同期工作以支持情境化的新颖性与贡献评估;2)隐式推理机制,深入理解复杂的研究动机与方法,并与同期相关工作进行全面比较,提升评估准确性。为引导推理过程,我们设计渐进式排名优化策略,促使LLM迭代改进预测结果,强调相对比较。在两个数据集上的实验表明,PaperEval在学术影响力与论文质量评估方面均优于现有方法。此外,我们将PaperEval部署于真实论文推荐系统中,用于筛选高质量论文,该系统在社交媒体上获得强烈关注——累计订阅数超8000人,多篇筛选论文播放量超过10000次,充分证明了其实际有效性。
原文摘要 · Abstract (English)
With the rapid and continuous increase in academic publications, identifying high-quality research has become an increasingly pressing challenge. While recent methods leveraging Large Language Models (LLMs) for automated paper evaluation have shown great promise, they are often constrained by outdated domain knowledge and limited reasoning capabilities. In this work, we present PaperEval, a novel LLM-based framework for automated paper evaluation that addresses these limitations through two key components: 1) a domain-aware paper retrieval module that retrieves relevant concurrent work to support contextualized assessments of novelty and contributions, and 2) a latent reasoning mechanism that enables deep understanding of complex motivations and methodologies, along with comprehensive comparison against concurrently related work, to support more accurate and reliable evaluation. To guide the reasoning process, we introduce a progressive ranking optimization strategy that encourages the LLM to iteratively refine its predictions with an emphasis on relative comparison. Experiments on two datasets demonstrate that PaperEval consistently outperforms existing methods in both academic impact and paper quality evaluation. In addition, we deploy PaperEval in a real-world paper recommendation system for filtering high-quality papers, which has gained strong engagement on social media -- amassing over 8,000 subscribers and attracting over 10,000 views for many filtered high-quality papers -- demonstrating the practical effectiveness of PaperEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。