检测大模型辅助审稿中的隐性偏见,发现机构排名影响最大
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- 通过操控作者信息,测试大模型审稿偏好
- 高排名机构作者获更高评分,边线论文易受偏见影响
- 模型对机构和资历有隐性偏好,适合关注公平性的研究者
大型语言模型(LLMs)正改变同行评审流程,从辅助撰写评价到自动生成完整评审。尽管带来新机遇,也引发公平性与可靠性担忧。本文通过控制作者元数据(包括机构、性别、资历和发表历史)的干预实验,系统分析了LLM生成评审中的偏见。结果一致显示,模型存在显著机构偏见,更倾向于高排名机构作者。同时,资历和过往发表记录也带来方向性偏好,可能影响临界论文的接受决策。性别影响较小但部分模型中仍存在。值得注意的是,在逐标记软评分层面,隐性偏见更为明显,表明对齐机制可能掩盖而非完全消除内在偏好。
原文摘要 · Abstract (English)
The adoption of large language models (LLMs) is transforming the peer review process, from assisting reviewers in writing detailed evaluations to generating entire reviews automatically. While these capabilities offer new opportunities, they also raise concerns about fairness and reliability. In this paper, we investigate bias in LLM-generated peer reviews through controlled interventions on author metadata, including affiliation, gender, seniority, and publication history. Our analysis consistently shows a strong affiliation bias favoring authors from highly ranked institutions. We also identify directional preferences associated with seniority and prior publication record, which can influence acceptance decisions for borderline papers. Gender effects are smaller but present in several models. Notably, implicit biases become more pronounced when examining token-level soft ratings, suggesting that alignment may mask but not fully eliminate underlying preferences
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。