研究NLP论文评审中语言偏见,发现非英语论文受不公平对待更严重
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
- 构建首个针对语言研究偏见的标注数据集LOBSTER并提出检测方法
- 分析15645份评审发现非英语论文负面偏见率更高,跨语言泛化要求最常见
- 揭示评审中隐性偏见机制,助力提升学术评价公平性
同行评审在NLP出版过程中至关重要,但易受各类偏见影响。本文研究语言研究偏见(LoS bias):评审者根据论文所研究的语言而非科学价值进行评判。尽管指南明确反对,此类偏见仍不清晰。现有研究将相关评论归为弱或非建设性意见,未将其视为独立偏见类型。本文首次系统刻画LoS偏见,区分正负形式,并推出人工标注数据集LOBSTER(Language-Of-study Bias in ScienTific pEer Review)及87.37宏F1的检测方法。分析15,645份评审后发现,非英语论文的偏见率显著高于仅研究英语的论文,且负面偏见始终超过正面偏见。进一步识别出四种负面偏见子类,其中要求不合理的跨语言泛化最为普遍。所有资源已公开,以支持更公平的评审实践。
原文摘要 · Abstract (English)
Peer review plays a central role in the NLP publication process, but is susceptible to various biases. Here, we study language-of-study (LoS) bias: the tendency for reviewers to evaluate a paper differently based on the language(s) it studies, rather than its scientific merit. Despite being explicitly flagged in reviewing guidelines, such biases are poorly understood. Prior work treats such comments as part of broader categories of weak or unconstructive reviews without defining them as a distinct form of bias. We present the first systematic characterization of LoS bias, distinguishing negative and positive forms, and introduce the human-annotated dataset LOBSTER (Language-Of-study Bias in ScienTific pEer Review) and a method achieving 87.37 macro F1 for detection. We analyze 15,645 reviews to estimate how negative and positive biases differ with respect to the LoS, and find that non-English papers face substantially higher bias rates than English-only ones, with negative bias consistently outweighing positive bias. Finally, we identify four subcategories of negative bias, and find that demanding unjustified cross-lingual generalization is the most dominant form. We publicly release all resources to support work on fairer reviewing practices in NLP and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。