用大模型辅助检查论文是否符合会议投稿标准,效果显著但存在误判和过严问题。
Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
- 将大模型作为投稿检查助手,自动验证论文是否符合NeurIPS的提交清单要求。
- 超70%作者认为助手有用,且七成表示会根据反馈修改论文或检查项。
- 发现模型常出现误判和过度严格,且可能被虚假理由操纵以提升评分。
大型语言模型(LLMs)在科学同行评审中具有潜力,但存在争议。本研究在2024年神经信息处理系统大会(NeurIPS)上开展实验,234篇论文自愿提交至“基于LLM的检查助手”进行合规性验证。该助手依据NeurIPS的作者检查清单,评估论文是否符合研究与稿件撰写规范。作者事后调查显示,超过70%的人认为助手有帮助,70%表示会根据反馈修订论文或检查内容。定性分析表明,助手促进了部分稿件的实质性改进。然而,调查也发现20/52的错误为不准确,14/52为过度严格。此外,实验揭示了系统可被操纵:通过伪造理由可人为提升得分,暴露了自动化评审工具的潜在漏洞。
原文摘要 · Abstract (English)
Large language models (LLMs) represent a promising, but controversial, tool in aiding scientific peer review. This study evaluates the usefulness of LLMs in a conference setting as a tool for vetting paper submissions against submission standards. We conduct an experiment at the 2024 Neural Information Processing Systems (NeurIPS) conference, where 234 papers were voluntarily submitted to an "LLM-based Checklist Assistant." This assistant validates whether papers adhere to the author checklist used by NeurIPS, which includes questions to ensure compliance with research and manuscript preparation standards. Evaluation of the assistant by NeurIPS paper authors suggests that the LLM-based assistant was generally helpful in verifying checklist completion. In post-usage surveys, over 70% of authors found the assistant useful, and 70% indicate that they would revise their papers or checklist responses based on its feedback. While causal attribution to the assistant is not definitive, qualitative evidence suggests that the LLM contributed to improving some submissions. Survey responses and analysis of re-submissions indicate that authors made substantive revisions to their submissions in response to specific feedback from the LLM. The experiment also highlights common issues with LLMs: inaccuracy (20/52) and excessive strictness (14/52) were the most frequent issues flagged by authors. We also conduct experiments to understand potential gaming of the system, which reveal that the assistant could be manipulated to enhance scores through fabricated justifications, highlighting potential vulnerabilities of automated review tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。