用专家共识测大模型教育观,发现它在争议领域有坚定立场。
How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models
- 用8个教育维度48项问题,量化模型偏好一致性
- 99.78%偏好传递性,92.79%准确率,基本符合专家共识
- 在情感与认知规范争议区,模型选择回应而非中立
本文首次系统测量大语言模型的教育适配度。采用经德尔菲法验证的8个教育理论维度共48项指标,研究发现GPT-5.1表现出高度一致的偏好模式(99.78%传递性;92.79%模型准确率),在人类专家共识存在的领域与人本主义教育原则高度契合。关键的是,当人类专家自身存在规范性分歧时,模型偏好也出现偏离,尤其在情感维度与认识论规范方面。这引发对对齐研究的根本疑问:当人类价值观本身存在争议时,模型应以何为对齐标准?研究结果表明,GPT-5.1在争议领域并非保持中立,而是形成连贯立场,更重视情感回应并拒绝虚假平衡。该方法结合德尔菲共识构建、结构化偏好诱导与瑟斯顿效用建模,为特定领域对齐评估提供可复现框架,超越通用价值基准。
原文摘要 · Abstract (English)
This paper presents the first systematic measurement of educational alignment in Large Language Models. Using a Delphi-validated instrument comprising 48 items across eight educational-theoretical dimensions, the study reveals that GPT-5.1 exhibits highly coherent preference patterns (99.78% transitivity; 92.79% model accuracy) that largely align with humanistic educational principles where expert consensus exists. Crucially, divergences from expert opinion occur precisely in domains of normative disagreement among human experts themselves, particularly emotional dimensions and epistemic normativity. This raises a fundamental question for alignment research: When human values are contested, what should models be aligned to? The findings demonstrate that GPT-5.1 does not remain neutral in contested domains but adopts coherent positions, prioritizing emotional responsiveness and rejecting false balance. The methodology, combining Delphi consensus-building with Structured Preference Elicitation and Thurstonian Utility modeling, provides a replicable framework for domain-specific alignment evaluation beyond generic value benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。