arXiv:2603.23659cs.CLcs.AI2026-03

探究大模型如何内化不同伦理框架,发现其表征存在差异但不完全独立。

Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges

  • 通过探针分析六款大模型在五种伦理框架下的内部表示
  • 不同伦理维度间存在不对称迁移,如义务论可部分泛化至美德伦理
  • 结果受模板表面特征影响,需谨慎解读,适合伦理计算研究者参考

当大型语言模型做出伦理判断时,其内部表征是否能区分不同的规范性框架,还是将伦理简化为单一可接受性维度?我们对六款参数规模从40亿到720亿的LLM,在五种伦理框架(义务论、功利主义、美德论、正义论、常识伦理)中进行了隐藏表征探测。分析揭示了具有差异化特征的伦理子空间,且存在非对称迁移模式——例如,义务论探针可部分泛化至美德情境,而常识探针在正义任务上表现严重崩溃。义务论与功利主义探针之间的分歧程度与模型行为熵正相关,但该关系可能部分源于对情景难度的共同敏感性。事后验证表明,探针结果部分依赖于基准模板的表面特征,提示需谨慎解释。本文讨论了方法带来的结构洞察及其认识论局限。

原文摘要 · Abstract (English)

When large language models make ethical judgments, do their internal representations distinguish between normative frameworks, or collapse ethics into a single acceptability dimension? We probe hidden representations across five ethical frameworks (deontology, utilitarianism, virtue, justice, commonsense) in six LLMs spanning 4B--72B parameters. Our analysis reveals differentiated ethical subspaces with asymmetric transfer patterns -- e.g., deontology probes partially generalize to virtue scenarios while commonsense probes fail catastrophically on justice. Disagreement between deontological and utilitarian probes correlates with higher behavioral entropy across architectures, though this relationship may partly reflect shared sensitivity to scenario difficulty. Post-hoc validation reveals that probes partially depend on surface features of benchmark templates, motivating cautious interpretation. We discuss both the structural insights these methods provide and their epistemological limitations.

伦理推理模型探针大模型表征认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。