用大模型检测两个百科全书的政治中立性,发现AI写的更不中立。
Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

- 用四个大模型对比分析1394篇政要条目在九个意识形态维度的中立性。
- 所有大模型都评出Grokipedia比Wikipedia更不中立,且倾向右翼经济立场。
- 研究揭示了大模型判断本身也存在意识形态模式,适合关注AI偏见的研究者阅读。
在线百科全书影响政治观点,并进而塑造民主讨论。2025年底,由大模型Grok完全撰写的《Grokipedia》问世,初衷是提供一个比常被指责有‘左翼’和‘自由主义’偏见的Wikipedia更无偏见的替代品。但由大模型撰写的百科是否真能实现更高中立性,还是只是嵌入了另一种意识形态?我们对Grokipedia和Wikipedia开展了大规模政治偏见研究,分析了1,394篇描述政府成员的文章,在九个专家标注的意识形态维度上,使用四名大模型裁判(Grok、Claude、Mistral、DeepSeek)进行评估。由于这些大模型自身也可能存在偏见,我们还考察了其评判模式。结果显示,所有大模型裁判(包括Grok)均认为Grokipedia的中立性低于Wikipedia。两本百科均整体呈现对政客的正面描绘,但偏向不同意识形态群体:Grokipedia特别倾向经济右翼政客,而惩罚社会自由派政客;Wikipedia则被评定为对社会自由派有明显偏好。
原文摘要 · Abstract (English)
Online encyclopedias shape political opinion and, through it, democratic discourse. In late 2025, Grokipedia was released, an encyclopedia written entirely by the LLM Grok. One motivation behind the project was to provide an unbiased alternative to Wikipedia, which has faced accusations of "left-wing" and "liberal" bias. But does an encyclopedia written by an LLM deliver greater neutrality, or does it simply embed a different ideology? We conduct a large-scale political bias study on Grokipedia and Wikipedia, analysing 1,394 article pairs describing members of government for neutrality along nine expert-coded ideology dimensions employing four LLM judges, Grok, Claude, Mistral, and DeepSeek. As the LLMs could themselves be biased, we also investigate patterns in their judgments. We find all LLM-judges, including Grok, to rate Grokipedia less neutral than Wikipedia. Both encyclopedias are rated as portraying politicians favourably overall, but towards different ideological groups. Grokipedia particularly favours economically right-wing politicians and penalises socially liberal ones, while Wikipedia is rated as favourably biased towards the latter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。