让人类先下判断,AI扩展成评论,提升审稿自动化与责任性平衡。
Judgment-Grounded Expansion for Peer Review Generation

- 人类给出评价结论,AI生成对应评论候选
- 用置信区间控制候选集大小,确保覆盖目标评论
- 适合需要可解释审稿的学术协作场景
自动审稿有助于加速科学发展。现有方法多采用端到端模式,但完全自动化难以满足问责需求。本文提出判断基扩展(judgment-grounded expansion)的人机协作范式:评审人提供评估主张,系统将其扩展为评论候选。我们将其建模为结构化生成-检查-优化流程,并通过用户研究收集了人机交互数据。针对可扩展评估和候选集筛选两个实际挑战,我们开发了大规模评估模拟方法,发现分位数预测(conformal prediction)能有效平衡候选数量与覆盖率。本工作确立了该任务的具体形式,并为未来协同审稿系统的设计提供了实证与方法基础。
原文摘要 · Abstract (English)
Automatic review generation is a promising direction for accelerating scientific progress. While most work adopts an end-to-end setup, its fully automated nature makes it less suitable for settings that demand accountability. To better balance automation and accountability, we formalize judgment-grounded expansion, a human-AI collaboration mode where a reviewer provides an evaluative claim and the system expands it into review comment candidate(s). We model it as a structured generate-check-refine process and conduct a user study to collect human-model interaction data. We study two practical challenges for judgment-grounded expansion: scalable evaluation and candidate set curation. We develop methods to simulate the process for large-scale evaluation, and show that conformal prediction is well suited to balancing candidate set size and target coverage. Our work establishes judgment-grounded expansion as a concrete task and provides empirical and methodological foundations for the design of future collaborative review generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。