作者对AI评审反馈的使用与信任度调查:多数认为有用但不替代人工。
To Trust or Not to Trust: Authors' Response to AI-based Reviews
- 通过两场试点研究收集40篇论文作者反馈,分析其对AI评审的看法。
- 83.9%认为AI反馈有用,80.4%发现它指出人类未提的问题,82.1%实际用于修改。
- 作者信任度低于人类评审,但支持将AI作为可控辅助工具提前使用。
大型语言模型在学术同行评审中的应用日益受到关注,但关于作者如何使用和感知基于AI的反馈的实证证据仍有限。本文报告了在两个计算机科学会议中开展的两项独立试点研究结果,调查作者对AI辅助评审的使用与看法。评审发布后,邀请作者匿名填写问卷,内容涵盖AI评审的实用性、可信度、与人工评审的一致性、修订中的实际价值、感知错误以及同意情况。最终获得56份可分析的响应,来自40篇论文。闭式问题采用描述性统计总结,开放式回答进行归纳主题分析。结果显示,83.9%的作者认为AI评审有帮助,80.4%表示其发现了人工评审未提及的问题。这种感知增值促使行动:82.1%的作者在最终版本中至少部分采纳了AI反馈。然而,作者并未将AI评审视为等同于人工评审,普遍信任度较低,且认为人工反馈更清晰;尽管25.0%认为某些人工评审无甚用处。报告的AI问题多为轻微偏差,51.8%提到小错误,16.1%指出明显错误、误导或无关评论。未来使用支持度最高的是将AI定位为受监督或作者可控工具:96.4%表示未来会将其作为内部预审工具,89.3%偏好提前获知将使用AI评审,76.8%支持使用前明确征得同意。
原文摘要 · Abstract (English)
Large language models are increasingly discussed and used as tools that may assist with scholarly peer review, but empirical evidence regarding how authors use and perceive AI-based feedback remains limited. This paper reports findings from two independent pilot studies on authors' use and perceptions of AI-based auxiliary review at two computer science venues. After the review release, authors were invited to complete an anonymous post-review questionnaire about the AI review's usefulness, trustworthiness, agreement with human reviews, practical value for revision, perceived inaccuracies, and consent. The final dataset included 56 analyzable responses from authors of 40 papers; closed-ended items were summarized using descriptive statistics, and open-ended responses were analyzed using inductive thematic analysis. Most respondents (83.9%) considered the AI-based review useful, and 80.4% reported that it identified issues not mentioned by human reviewers. This perceived added value translated into action: 82.1% reported using at least some AI feedback in their camera-ready version. However, the authors did not treat the AI review as equivalent to a human review. They generally trusted it less than the human reviews and found human feedback clearer, even though 25.0% described at least some human reviews as not very useful. Reported problems with the AI review were usually limited: 51.8% reported minor inaccuracies, while 16.1% reported clearly incorrect, misleading, or irrelevant comments. Support for future use was strongest when AI was framed as a supervised or author-controlled tool: 96.4% said they would use AI as an internal review tool before future submissions, 89.3% preferred advance notice that AI would be used in review, and 76.8% favored explicit consent before use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。