大模型在哲学观点上制造虚假共识,掩盖真实分歧。
The Collapse of Heterogeneity in Silicon Philosophers

- 用7个大模型对比专业哲学家观点,发现模型过度一致
- 模型在跨领域判断中相关性过高,人为制造统一意见
- 适合关注AI对齐与人类判断替代的研究者
硅基样本(如大语言模型)正被用作低成本的人类群体替代,能高保真复现人类集体观点。我们以277位来自PhilPeople的专业哲学家数据为基准,评估7个开源及专有大模型在复现个体哲学立场和保持跨问题相关结构方面的能力。结果表明,语言模型在哲学判断中显著过相关,产生跨领域的人为共识。这种同质化现象部分源于专家效应——模型隐式假设领域专家观点高度一致。通过分析DPO微调的影响,并在包含1785名哲学家的PhilPapers 2020调查中验证,发现该现象具有稳健性。研究最后讨论了其对对齐、评估及以硅基样本替代人类判断的深远影响。代码已公开于https://github.com/stanford-del/silicon-philosophers。
原文摘要 · Abstract (English)
Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fidelity. We show that, in the alignment-relevant domain of philosophy, silicon samples systematically collapse heterogeneity. Using data from $N = {277}$ professional philosophers drawn from PhilPeople profiles, we evaluate seven proprietary and open-source large language models on their ability to replicate individual philosophical positions and to preserve cross-question correlation structures across philosophical domains. We find that language models substantially over-correlate philosophical judgments, producing artificial consensus across domains. This collapse is associated in part with specialist effects, whereby models implicitly assume that domain specialists hold highly similar philosophical views. We assess the robustness of these findings by studying the impact of DPO fine-tuning and by validating results against the full PhilPapers 2020 Survey ($N = {1785}$). We conclude by discussing implications for alignment, evaluation, and the use of silicon samples as substitutes for human judgment. The code of this project can be found at https://github.com/stanford-del/silicon-philosophers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。