用大模型揭示科学评审中隐藏的潜规则,让隐性标准显性化。
Language Models Should be Used to Surface the Unwritten Code of Science and Society
- 让大模型通过对比论文生成自我解释,挖掘评审背后的隐性标准。
- 发现评审偏爱叙事连接而非理论严谨,但不公开承认这一点。
- 为检测社会隐性规范提供可复用框架,适合政策与AI伦理研究者。
本文呼吁研究者不仅关注大语言模型如何继承人类偏见,更应利用这些偏见揭示社会与科学中的'未言明规则'——如隐性刻板印象与启发式判断。我们以学术同行评审为例,提出一个概念框架:通过让大模型在46个会议提交的成对论文中生成自洽假设(为何某篇得分更高),并迭代探索无法解释的配对,揭示其推理逻辑。结果发现,模型的初始规范先验(如理论严谨性)逐渐被后验更新为强调外部关联(如文献定位与跨领域联系)。人类审稿人虽在评分上隐性奖励此类叙事特征(相关性= -0.14),却极少在评语中提及;而对符合模型先验的方面则显性认可(相关性=0.49)。该模式在不同模型和跨样本判断中均稳健。本框架可推广至各类社会隐性规范的诊断,推动公众讨论与负责任AI发展。
原文摘要 · Abstract (English)
This paper calls on the research community not only to investigate how human biases are inherited by large language models (LLMs) but also to explore how these biases in LLMs can be leveraged to make society's "unwritten code" - such as implicit stereotypes and heuristics - visible and accessible for critique. We introduce a conceptual framework through a case study in science: uncovering hidden rules in peer review - the factors that reviewers care about but rarely state explicitly due to normative scientific expectations. The idea of the framework is to push LLMs to speak out their heuristics through generating self-consistent hypotheses - why one paper appeared stronger in reviewer scoring - among paired papers submitted to 46 academic conferences, while iteratively searching deeper hypotheses from remaining pairs where existing hypotheses cannot explain. We observed that LLMs' normative priors about the internal characteristics of good science extracted from their self-talk, e.g., theoretical rigor, were systematically updated toward posteriors that emphasize storytelling about external connections, such as how the work is positioned and connected within and across literatures. Human reviewers tend to explicitly reward aspects that moderately align with LLMs' normative priors (correlation = 0.49) but avoid articulating contextualization and storytelling posteriors in their review comments (correlation = -0.14), despite giving implicit reward to them with positive scores. These patterns are robust across different models and out-of-sample judgments. We discuss the broad applicability of our proposed framework, leveraging LLMs as diagnostic tools to amplify and surface the tacit codes underlying human society, enabling public discussion of revealed values and more precisely targeted responsible AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。