NAIPv2通过配对学习提升论文质量评估的准确性与效率。
NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation
- 在领域-年份分组内采用配对学习,减少评分偏差。
- 达78.2% AUC与0.432斯皮尔曼相关系数,推理高效线性扩展。
- 适合需要高泛化能力的论文评审自动化场景。
科学论文质量评估对人类和人工智能推动科学发展至关重要。现有基于大语言模型的方法推理成本高,而快速的直接评分回归法受限于尺度不一致。本文提出NAIPv2,一种去偏且高效的论文质量评估框架。NAIPv2在领域-年份分组内采用配对学习以降低审稿人评分的不一致性,并引入评审倾向信号(RTS),作为评分与置信度的概率集成。为支持训练与评估,我们构建了包含24,276篇ICLR投稿的NAIDv2大规模数据集,附带元数据与结构化内容。模型基于配对比较训练,但部署时可实现高效的点式预测,达到78.2% AUC与0.432斯皮尔曼相关系数的领先性能,同时保持线性时间推理效率。值得注意的是,在未见的NeurIPS投稿上,其预测得分在从拒稿到口头报告的决策类别中持续上升,展现出强泛化能力。这些结果确立了NAIPv2作为去偏、可扩展的自动化论文质量评估框架,迈向未来科学智能系统的重要一步。代码与数据集已发布于 sway.cloud.microsoft/Pr42npP80MfPhvj8。
原文摘要 · Abstract (English)
The ability to estimate the quality of scientific papers is central to how both humans and AI systems will advance scientific knowledge in the future. However, existing LLM-based estimation methods suffer from high inference cost, whereas the faster direct score regression approach is limited by scale inconsistencies. We present NAIPv2, a debiased and efficient framework for paper quality estimation. NAIPv2 employs pairwise learning within domain-year groups to reduce inconsistencies in reviewer ratings and introduces the Review Tendency Signal (RTS) as a probabilistic integration of reviewer scores and confidences. To support training and evaluation, we further construct NAIDv2, a large-scale dataset of 24,276 ICLR submissions enriched with metadata and detailed structured content. Trained on pairwise comparisons but enabling efficient pointwise prediction at deployment, NAIPv2 achieves state-of-the-art performance (78.2% AUC, 0.432 Spearman), while maintaining scalable, linear-time efficiency at inference. Notably, on unseen NeurIPS submissions, it further demonstrates strong generalization, with predicted scores increasing consistently across decision categories from Rejected to Oral. These findings establish NAIPv2 as a debiased and scalable framework for automated paper quality estimation, marking a step toward future scientific intelligence systems. Code and dataset are released at sway.cloud.microsoft/Pr42npP80MfPhvj8.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。