提出新指标评估论文评审中论点与证据的逻辑关联,提升评审自动化效率。
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
- 通过分析论点与证据间的逻辑推理关系评估论证强度。
- 在人类评分相关性上优于传统方法,相关系数提升显著。
- 适合用于自动化评审系统,辅助科研人员高效处理审稿任务。
由于投稿论文数量激增,科学同行评审面临人力短缺问题。利用语言模型降低人工成本成为潜在解决方案。本文提出一种可解释的科学评审评价方法,通过提取论证的核心成分——论点与证据,并基于论点被证据支持的比例来评估论证充分性。论证充分性指论点基于客观事实的程度。然而,仅判断证据是否存在不足以准确评估论证质量,还需精确评估论点与证据之间的逻辑推断关系。为此,本文提出一种新的评审评论评估指标,专门衡量论点与证据间的逻辑推理。实验结果表明,该方法与人类评分的相关性高于传统方法,展现出提升同行评审效率的潜力。
原文摘要 · Abstract (English)
The scientific peer-review process is facing a shortage of human resources due to the rapid growth in the number of submitted papers. The use of language models to reduce the human cost of peer review has been actively explored as a potential solution to this challenge. A method has been proposed to evaluate the level of substantiation in scientific reviews in a manner that is interpretable by humans. This method extracts the core components of an argument, claims and evidence, and assesses the level of substantiation based on the proportion of claims supported by evidence. The level of substantiation refers to the extent to which claims are based on objective facts. However, when assessing the level of substantiation, simply detecting the presence or absence of supporting evidence for a claim is insufficient; it is also necessary to accurately assess the logical inference between a claim and its evidence. We propose a new evaluation metric for scientific review comments that assesses the logical inference between claims and evidence. Experimental results show that the proposed method achieves a higher correlation with human scores than conventional methods, indicating its potential to better support the efficiency of the peer-review process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。