arXiv:2603.20214cs.CYcs.AI2026-03综述

AI辅助学术评审需明确人机分工,避免责任模糊与认知伤害。

Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge

  • 通过社交媒体与访谈分析,发现AI仅适合辅助修改表述、组织反馈。
  • 核心判断如创新性、贡献度必须由人类负责,避免标准化过度与责任不清。
  • 建议制定角色化管控规则,保障评审公平与可问责性,尤其保护青年学者。

生成式AI正逐步进入学术同行评审流程,引发公平性、问责制与评价正当性的关切。尽管此类系统在审稿人负担加重的背景下承诺提升效率,却带来新的社会技术风险。本文结合对448条社交媒体讨论内容的语篇分析及14位顶级人工智能与人机交互会议领域主席和程序主席的深度访谈,探究生成式AI在同行评审中的讨论与实际体验。两个数据集均显示,广泛认同生成式AI可用于有限支持任务,如提升语言清晰度或结构化反馈,但核心评价判断——包括新颖性、贡献度与录用决定——应仍由人类承担。同时,受访者也指出存在认识论伤害、过度标准化、责任不明以及提示注入等对抗性风险。访谈揭示结构性压力与制度政策模糊性将解释与执行负担转嫁给个体学者,尤以初级作者和审稿人为甚。通过公共治理话语与真实评审实践的三角验证,本研究将AI介入的同行评审重构为一项社会技术治理挑战,并提出维护问责、信任与实质性人类监督的建议。总体认为,不应采取一刀切禁令或仅依赖检测,而应明确保留评价判断的人类主权,同时建立可执行的角色特定控制机制。最后提出具体角色导向的规范建议,以界定支持与判断的边界。

原文摘要 · Abstract (English)

Generative AI tools are increasingly entering academic peer review workflows, raising questions about fairness, accountability, and the legitimacy of evaluative judgment. While these systems promise efficiency gains amid growing reviewer overload, their use introduces new sociotechnical risks. This paper presents a convergent mixed-method study combining discourse analysis of 448 social media posts with interviews with 14 area chairs and program chairs from leading AI and HCI conferences to examine how GenAI is discussed and experienced in peer review. Across both datasets, we find broad agreement that GenAI may be acceptable for limited supportive tasks, such as improving clarity or structuring feedback, but that core evaluative judgments, assessing novelty, contribution, and acceptance, should remain human responsibilities. At the same time, participants highlight concerns about epistemic harm, over-standardization, unclear responsibility, and adversarial risks such as prompt injection. User interviews reveal how structural strain and institutional policy ambiguity shift interpretive and enforcement burdens onto individual scholars, disproportionately affecting junior authors and reviewers. By triangulating public governance discourse with lived review practices, this work reframes AI mediated peer review as a sociotechnical governance challenge and offers recommendations for preserving accountability, trust, and meaningful human oversight. Overall, we argue that AI-assisted peer review is best governed not by blanket bans or detection alone, but by explicitly reserving evaluative judgment for humans while instituting enforceable, role-specific controls that preserve accountability. We conclude with role specific recommendations that formalize the support judgment boundary.

AI治理同行评审人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。