arXiv:2605.26662cs.CLcs.AI2026-05

评估论文中AI使用率时,忽略领域和国家差异会导致严重误判。

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

论文配图:AI evaluation may bias perceptions: The importance of context in interpreting academic writing
图 1 · 摘自论文原文
  • 按国家和领域分别构建文本相似性基准,避免风格差异干扰
  • 统一基准使部分国家/领域误判率超30%,特定领域偏差达25%以上
  • 适合关注AI评估公平性的科研管理者与期刊编辑

本文研究了在不同国家和学科背景下,若评估方法忽略上下文差异,会对科学写作中AI使用率的估算产生系统性偏差。基于Dimensions数据库的大规模期刊论文数据,我们构建了基于人类撰写与大模型重写摘要差异的AI相似度基准。结果表明,合并使用的统一基准会混淆原本存在的写作风格差异,导致即使在大模型出现前的文献中也产生显著扭曲,尤其在不同国家-学科组间偏差明显。相比之下,按国家-学科定制的基准能有效缓解此类偏差,提供更可信的参照标准。将该方法应用于2025年发表的文献分析显示,统一基准系统性高估某些国家与学科的AI使用率,而低估其他群体。研究强调,实现对科学领域中AI使用的准确、公平评估,必须采用上下文感知的测量方法。

原文摘要 · Abstract (English)

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.

AI评估学术写作偏见校正跨域比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。