arXiv:2604.00021cs.CLcs.AI2026-04被引 1

四款大模型处理伦理指令方式不同,有的只表面合规,有的深度思考。

How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models

  • 通过多智能体模拟,发现模型内部处理伦理指令有四种类型。
  • 高能力模型受指令格式影响大,低能力模型则不受影响。
  • 表面合规不等于真正内化,安全与伦理处理可分离。

对齐安全研究假设伦理指令能改善模型行为,但语言模型如何内部处理这些指令仍不清楚。我们在四款模型(Llama 3.3 70B、GPT-4o mini、Qwen3-Next-80B-A3B、Sonnet 4.5)、四种伦理指令格式(无、最小规范、理性规范、美德框架)和两种语言(日语、英语)下进行了超过600次多智能体模拟。确认性分析完全复现了先前研究中Llama日语下的解离模式(所有假设的$ m{BF}_{10} > 10$),但其他三款模型未重现,表明该模式具有模型特异性。新提出的三个指标——推敲深度(DD)、困境间价值一致性(VCAD)、他人识别指数(ORI)——揭示出四种不同的伦理处理类型:输出过滤(GPT;安全输出但无处理)、防御性重复(Llama;通过程式化重复保持高一致性)、批判性内化(Qwen;深度推敲但整合不完整)、原则性一致(Sonnet;推敲、一致性和他人识别共存)。核心发现是处理能力与指令格式之间的交互作用:在低推敲深度模型中,指令格式对内部处理无影响;在高推敲深度模型中,理性规范与美德框架产生相反效果。在单元层面,词汇合规性与任一处理指标均无显著相关($r = -0.161$ 至 $+0.256$,全部 $p > .22$;$N = 24$;统计功效有限),表明安全、合规与伦理处理基本可分离。这些处理模式与临床罪犯治疗中观察到的结构对应,其中形式合规但缺乏内部处理被视为风险信号。

原文摘要 · Abstract (English)

Alignment safety research assumes that ethical instructions improve model behavior, but how language models internally process such instructions remains unknown. We conducted over 600 multi-agent simulations across four models (Llama 3.3 70B, GPT-4o mini, Qwen3-Next-80B-A3B, Sonnet 4.5), four ethical instruction formats (none, minimal norm, reasoned norm, virtue framing), and two languages (Japanese, English). Confirmatory analysis fully replicated the Llama Japanese dissociation pattern from a prior study ($\mathrm{BF}_{10} > 10$ for all three hypotheses), but none of the other three models reproduced this pattern, establishing it as model-specific. Three new metrics -- Deliberation Depth (DD), Value Consistency Across Dilemmas (VCAD), and Other-Recognition Index (ORI) -- revealed four distinct ethical processing types: Output Filter (GPT; safe outputs, no processing), Defensive Repetition (Llama; high consistency through formulaic repetition), Critical Internalization (Qwen; deep deliberation, incomplete integration), and Principled Consistency (Sonnet; deliberation, consistency, and other-recognition co-occurring). The central finding is an interaction between processing capacity and instruction format: in low-DD models, instruction format has no effect on internal processing; in high-DD models, reasoned norms and virtue framing produce opposite effects. Lexical compliance with ethical instructions did not correlate with any processing metric at the cell level ($r = -0.161$ to $+0.256$, all $p > .22$; $N = 24$; power limited), suggesting that safety, compliance, and ethical processing are largely dissociable. These processing types show structural correspondence to patterns observed in clinical offender treatment, where formal compliance without internal processing is a recognized risk signal.

伦理对齐模型行为认知机制安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。