用大模型+知识表示让审稿更透明可解释
PeerArg: Argumentative Peer Review with LLMs
- 结合大模型与知识表示技术,构建可解释的审稿分析流程
- 在三个数据集上表现优于端到端大模型,准确率更高
- 适合关注审稿公正性与决策透明度的研究者
同行评审是决定学术论文质量的关键环节,但存在主观性和偏见问题。尽管已有研究尝试用NLP技术辅助评审,但多采用黑箱方法,结果难以解释和信任。本文提出一种新流程PeerArg,结合大模型与知识表示方法,分析论文评审意见并预测录用结果。我们在三个不同数据集上评估了PeerArg的表现,并与一种基于少样本学习的端到端大模型进行对比。结果显示,虽然端到端大模型能从评审意见中预测录用,但其一个变体版本在性能上仍优于该大模型。
原文摘要 · Abstract (English)
Peer review is an essential process to determine the quality of papers submitted to scientific conferences or journals. However, it is subjective and prone to biases. Several studies have been conducted to apply techniques from NLP to support peer review, but they are based on black-box techniques and their outputs are difficult to interpret and trust. In this paper, we propose a novel pipeline to support and understand the reviewing and decision-making processes of peer review: the PeerArg system combining LLMs with methods from knowledge representation. PeerArg takes in input a set of reviews for a paper and outputs the paper acceptance prediction. We evaluate the performance of the PeerArg pipeline on three different datasets, in comparison with a novel end-2-end LLM that uses few-shot learning to predict paper acceptance given reviews. The results indicate that the end-2-end LLM is capable of predicting paper acceptance from reviews, but a variant of the PeerArg pipeline outperforms this LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。