用真实审稿记录自动验证论文修改是否真正回应了意见。
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

- 基于审稿记录构建可解释的验证机制,追踪修改是否匹配反馈
- 最佳模型仅达0.501分,证明证据验证仍是核心难题
- 适合关注科学流程可信度与AI辅助审稿的研究者
大型语言模型(LLMs)已能协助科学研究与同行评审,但一个关键能力仍待深入:验证审稿意见是否带来真实、有证据支持的稿件改进。我们提出AutoSupervision,通过可追溯的审稿记录自动评估修订是否真正回应了审稿人关切。该系统利用透明的同行评审记录作为自然监督信号——审稿意见指出科学疑点,作者回复描述解决措施,修订稿件提供证据。给定审稿意见、作者回复和修订稿,模型需识别审稿关切、判断是否解决,并定位支撑性文本证据。我们基于56,000篇Nature Communications文章及其对应的评审记录构建了AutoSupervision数据集。实验显示,尽管大模型在识别审稿关切方面表现良好(如GPT-5.5得分为0.754),但基于证据的验证仍是主要瓶颈,最优模型得分仅为0.501。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。