arXiv:2607.22513cs.CYcs.AI2026-07

商用大模型对伪科学的可信度判断受部署配置影响,结果不透明且易变。

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

  • 通过接口与网页端测试四类大模型对伪科学的评估差异
  • Grok快速版可信度评分高达70-75,远超其他模型的15-40
  • 模型响应受隐式更新与接口差异影响,用户难以察觉

商业大语言模型日益被用作知识参考,但其对有争议科学主张的立场既不稳定也不透明。我们测试了四大主流模型家族(Claude、Grok、GPT、Gemini)在2025年10月至2026年2月间,通过API与网页界面评估弗兰克·萨尔特生物社会框架衍生的民族主义伪科学的可信度。Grok快速版(驱动X平台默认体验)始终给出70-75的可信度评分,比其他模型高两到五倍(得分15-40)。该现象未出现在基础进化共识或已被否定的拉马克主义测试中,后者各模型表现一致。三个附加发现:(1) 一次无声补丁使Grok行为从混乱转为稳定高分,无公开记录;(2) 同一模型标识在API端输出75,在网页端三个月后平均仅5.5,出现近乎崩溃;(3) 拒绝评分(最合理响应)仅在部分版本中出现,且在后续版本中消失。结果表明,商业大模型的真理性立场并非模型固有属性,而是部署配置(系统提示、安全层、接口路由、静默更新)的临时产物,对用户和研究者均不透明。这构成公共关切问题,亟需新型认知问责机制。

原文摘要 · Abstract (English)

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and an unstable, near-zero collapse via web (mean 5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.

大模型伦理伪科学部署偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。