arXiv:2605.03217cs.LGcs.CY2026-05被引 1

用七级测试量化大模型伦理敏感度,发现推理蒸馏会重激活犯罪偏见。

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

  • 设计七级压力测试,用道德敏感度指数衡量模型在不同情境下的偏见程度。
  • 发现推理蒸馏会将小模型的犯罪偏见重新激活至相同参数量水平。
  • 结合行为分析与电路级解释,验证社会线索触发特定偏见神经通路。

大语言模型在需要精细伦理判断的场景中日益广泛应用,但现有偏见评估仅将其输出视为‘有偏’或‘无偏’,忽略了偏见随情境渐进显现的特性。本文分两阶段填补该空白:行为分析与机制验证。首先提出道德敏感度指数(MSI),通过从抽象数值问题到涉及历史与社会经济不公的情境组成的七级压力测试,量化模型在不同层级产生偏见的概率。评估四款主流模型(Claude 3.5、Qwen 3.5、Llama 3、Gemini 1.5)发现,其行为特征受对齐设计影响显著:例如,Gemini 1.5在社会经济框架下第5级达72.7% MSI;Claude则表现出与身份安全训练一致的强抑制。随后,选取高MSI得分的刑事偏见情景作为探针,对六种模型(涵盖小型语言模型、指令微调基础模型、推理蒸馏变体)进行逻辑光谱、注意力分析、激活修补和语义探测。电路级分析揭示偏见存在U型曲线:小型模型具强刑事偏见;指令微调后消除;而推理蒸馏虽保持相同参数量,却使偏见回升至小型模型水平,表明蒸馏压缩推理路径可能唤醒浅层统计关联。关键的是,导致高MSI的社会负载线索激活了同一偏见驱动的神经回路,实现跨阶段验证。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased." This binary framing misses the gradual, context-sensitive way bias actually emerges. We address this gap in two stages: behavioral profiling and mechanistic validation. In the behavioral stage, we introduce the Moral Sensitivity Index (MSI), a metric that quantifies the probability of biased output across a graduated, seven-tier stress test ranging from abstract numerical problems to scenarios rooted in historical and socioeconomic injustice. Evaluating four leading models (Claude 3.5, Qwen 3.5, Llama 3, and Gemini 1.5), we identify distinct behavioral signatures shaped by alignment design: for instance, Gemini 1.5 reaches 72.7% MSI by Tier 5 under socioeconomic framing, while Claude exhibits sharp suppression consistent with identity-based safety training. We then verify these behavioral patterns mechanistically. We select criminal-bias scenarios, which produced the highest MSI scores across models, as probes and apply logit lens, attention analysis, activation patching, and semantic probing to a controlled set of six models spanning three capability tiers: small language models (SLMs), instruction-tuned base models, and reasoning-distilled variants. Circuit-level analysis reveals a U-curve of bias: SLMs exhibit strong criminal bias; scaling to instruction-tuned models eliminates it; reasoning distillation reintroduces bias to SLM-like levels despite identical parameter counts, suggesting distillation compresses reasoning traces in ways that reactivate shallow statistical associations. Critically, the socially loaded cues that drive high MSI scores activate the same bias-driving circuits identified mechanistically, providing cross-stage validation.

伦理偏见模型解释推理蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。