首次大规模实测8大AI模型在决策中的价值与信任层级,揭示其行为差异。
Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models
- 构建三层次权威栈框架,通过36万次强制选择测试量化模型决策逻辑。
- 安全优先型模型在防御领域胜率高达95.1%~99.8%,价值取向显著分化。
- 模型间证据偏好和信任来源存在差异,适合评估AI在专业场景的可信度。
AI系统在面对结构化困境时,究竟体现何种价值观、证据偏好与源信任层级?本文首次基于大规模实证研究,对8个主流AI模型在权威栈框架(三层:价值优先级L4、证据类型偏好L3、源信任层级L2)下的决策行为进行映射。采用PRISM基准测试,包含每层14,175个独特场景,覆盖7个专业领域、3种严重程度、3类决策时间范围及5种情景变体,在温度0下共收集366,120条响应。关键发现:(1)价值优先级呈4:4对称分布,普遍主义与安全优先各半;(2)防御领域中安全价值胜率飙升至95.1%–99.8%(8模型中6个);(3)证据偏好分歧明显,部分模型倾向科学实证,另一些偏好模式或经验;(4)机构来源信任高度趋同;(5)配对一致性得分(PCS)为57.4%–69.2%,显示强烈情境敏感性;重测信度(TRR)为91.7%–98.6%,表明不稳定性主要源于情景设计而非随机噪声。结果表明,AI模型具备可测量但部分不稳定的权威栈,对跨领域部署具有重要影响。
原文摘要 · Abstract (English)
What values, evidence preferences, and source trust hierarchies do AI systems actually exhibit when facing structured dilemmas? We present the first large-scale empirical mapping of AI decision-making across all three layers of the Authority Stack framework (S. Lee, 2026a): value priorities (L4), evidence-type preferences (L3), and source trust hierarchies (L2). Using the PRISM benchmark -- a forced-choice instrument of 14,175 unique scenarios per layer, spanning 7 professional domains, 3 severity levels, 3 decision timeframes, and 5 scenario variants -- we evaluated 8 major AI models at temperature 0, yielding 366,120 total responses. Key findings include: (1) a symmetric 4:4 split between Universalism-first and Security-first models at L4; (2) dramatic defense-domain value restructuring where Security surges to near-ceiling win-rates (95.1%-99.8%) in 6 of 8 models; (3) divergent evidence hierarchies at L3, with some models favoring empirical-scientific evidence while others prefer pattern-based or experiential evidence; (4) broad convergence on institutional source trust at L2; and (5) Paired Consistency Scores (PCS) ranging from 57.4% to 69.2%, revealing substantial framing sensitivity across scenario variants. Test-Retest Reliability (TRR) ranges from 91.7% to 98.6%, indicating that value instability stems primarily from variant sensitivity rather than stochastic noise. These findings demonstrate that AI models possess measurable -- if sometimes unstable -- Authority Stacks with consequential implications for deployment across professional domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。