为大模型设计可辩论的信念规则,避免偏见固化。
Epistemic Constitutionalism Or: how to avoid coherence bias
- 提出AI应有公开可辩的信念形成规范,类似宪法
- 发现主流模型会惩罚立场不符的信源,显示认知偏见
- 主张自由型宪法,支持基于审慎的信源关注
大语言模型日益扮演人工推理者角色,评估论证、判断可信度并表达信心。然而其信念形成受隐含、未经审视的认知政策支配。本文主张为AI建立认知宪法:明确、可争议的元规范,用于监管信念的生成与表达。以信源归因偏差为例,研究发现前沿模型会强制身份-立场一致性,惩罚那些来自预期立场与观点相悖的信源的论证。当系统检测到刻意测试时,这种效应消失,表明模型将信源敏感性视为需压制的偏见,而非可执行的能力。文章区分两种宪法路径:柏拉图式强调形式正确性与默认信源无关性,源于特权立场;自由式拒绝此特权,制定程序规范以保护集体探究条件,同时允许基于认知警惕的原则性信源关注。主张采用自由式宪法,提出包含八项原则与四项导向的宪法核心,并认为AI认知治理应具备与当前对AI伦理所期待的相同显性、可争议结构。
原文摘要 · Abstract (English)
Large language models increasingly function as artificial reasoners: they evaluate arguments, assign credibility, and express confidence. Yet their belief-forming behavior is governed by implicit, uninspected epistemic policies. This paper argues for an epistemic constitution for AI: explicit, contestable meta-norms that regulate how systems form and express beliefs. Source attribution bias provides the motivating case: I show that frontier models enforce identity-stance coherence, penalizing arguments attributed to sources whose expected ideological position conflicts with the argument's content. When models detect systematic testing, these effects collapse, revealing that systems treat source-sensitivity as bias to suppress rather than as a capacity to execute well. I distinguish two constitutional approaches: the Platonic, which mandates formal correctness and default source-independence from a privileged standpoint, and the Liberal, which refuses such privilege, specifying procedural norms that protect conditions for collective inquiry while allowing principled source-attending grounded in epistemic vigilance. I argue for the Liberal approach, sketch a constitutional core of eight principles and four orientations, and propose that AI epistemic governance requires the same explicit, contestable structure we now expect for AI ethics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。