用大模型提前模拟审稿意见,帮作者发现潜在问题。
More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
- 构建大模型系统,在投稿前生成大量原子化审稿关切点。
- 压缩后报告覆盖84.9%重要问题,但需5.2倍更多计算量。
- 适合希望提前优化论文的作者,尤其关注审稿反馈质量者。
审稿反馈常过晚,作者无法有效修改。本文研究一种面向作者的大模型系统,将部分审稿压力前置:自动生成大量原子级问题,并压缩为简短报告。基于10,000篇ICLR 2026投稿中3,398篇可获取的预审稿版本,评估其与历史审稿的一致性及遗漏问题的有效性。在10篇诊断论文上,独立采样覆盖44.9%历史问题;经去重与补全后,严格覆盖率达78.7%,严重性加权覆盖达84.9%,但需3.6倍请求和5.2倍令牌开销。隐藏的Top-32最优选择器可保留79.3%加权覆盖率(来自256候选池),而仅基于论文的选择器仅保留40–44%。结果表明,大模型能提供广泛覆盖,但压缩能力不足,消融实验指出代表性筛选与匹配敏感性是主要瓶颈。
原文摘要 · Abstract (English)
Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separately, the possible validity of concerns they omit. From 10,000 ICLR 2026 submissions, we use 3,398 manuscripts with accessible versions that predate review. On a ten-paper diagnostic, independent sampling covers 44.9% of historical issues; deduplication and refill reaches 78.7% strict and 84.9% seriousness-weighted coverage, at 3.6$\times$ more requests and 5.2$\times$ more tokens. A hidden Top-32 Oracle preserves the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retain only 40--44%. LLM review therefore provides broad coverage with a large candidate pool but compresses poorly; ablations identify representative selection and matcher sensitivity as the main sources of this gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。