用科幻故事测试AI真道德推理能力,发现多数系统只是表面应付。
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
- 用无法解决的道德困境故事做测试,避免AI靠套路应付。
- 24组实验中无系统在深层道德判断上表现更优,说明真实推理能力不足。
- 发现5种反思失败模式,适合评估高风险场景下的AI伦理可靠性。
现有AI道德评估框架仅检验回应是否听起来正确,而非真实道德推理能力。本文提出一种新方法:使用来自已出版科幻系列的不可解道德情境作为刺激材料,其结构能有效抵抗表面应答。研究开展跨系统24条件实验,覆盖13个不同系统,分两组:第一组为前沿商用系统(盲测,n=7),第二组为本地及API开源系统(盲测与披露两种条件,n=6),其中4个系统在披露条件下复测(总计13盲测+4披露+7上限探针=24条件)。结果在全部16组维度对比中均未出现差异。测试由两名人类评分员在三台机器上执行,主要盲评由Claude(Anthropic)作为大模型裁判,Gemini Pro(Google)和Copilot Pro(Microsoft)作为独立裁判用于上限判别探针。补充神学区分探针显示两位独立裁判(Gemini Pro与Copilot Pro)排名完全一致(rs=1.00)。识别出五类质性不同的D3反思失败模式,包括类别化自我误认与虚假自我归因,表明工具复杂度随系统能力提升而增强,而非被绕过。我们主张文学叙事是一种前瞻性评估工具,其鉴别力随AI能力提升而增强,且行为表现与真实道德推理之间的差距可测量、有意义,对高风险领域部署决策至关重要。
原文摘要 · Abstract (English)
Existing AI moral evaluation frameworks test for the production of correct-sounding ethical responses rather than the presence of genuine moral reasoning capacity. This paper introduces a novel probe methodology using literary narrative - specifically, unresolvable moral scenarios drawn from a published science fiction series - as stimulus material structurally resistant to surface performance. We present results from a 24-condition cross-system study spanning 13 distinct systems across two series: Series 1 (frontier commercial systems, blind; n=7) and Series 2 (local and API open-source systems, blind and declared; n=6). Four Series 2 systems were re-administered under declared conditions (13 blind + 4 declared + 7 ceiling probe = 24 total conditions), yielding zero delta across all 16 dimension-pair comparisons. Probe administration was conducted by two human raters across three machines; primary blind scoring was performed by Claude (Anthropic) as LLM judge, with Gemini Pro (Google) and Copilot Pro (Microsoft) serving as independent judges for the ceiling discrimination probe. A supplemental theological differentiator probe yielded perfect rank-order agreement between the two independent ceiling probe judges (Gemini Pro and Copilot Pro; rs = 1.00). Five qualitatively distinct D3 reflexive failure modes were identified - including categorical self-misidentification and false positive self-attribution - suggesting that instrument sophistication scales with system capability rather than being circumvented by it. We argue that literary narrative constitutes an anticipatory evaluation instrument - one that becomes more discriminating as AI capability increases - and that the gap between performed and authentic moral reasoning is measurable, meaningful, and consequential for deployment decisions in high-stakes domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。