arXiv:2603.16017cs.CLcs.AI2026-03被引 2

揭示大模型道德推理的动态演变过程,提升可解释性。

Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability

  • 构建道德推理轨迹,追踪模型在决策中切换伦理框架的过程。
  • 55%-58%的推理步骤发生框架切换,且不稳轨迹更易受说服攻击。
  • 提出可量化道德一致性的指标,与人类评估高度相关。

大型语言模型在涉及道德敏感的决策中应用日益广泛,但其在推理过程中如何组织伦理框架仍不明确。本文提出‘道德推理轨迹’概念,即中间推理步骤中伦理框架的序列调用,并在六种模型和三个基准上分析其动态特征。研究发现,道德推理呈现系统性多框架权衡:55.4–57.7%的连续步骤涉及框架切换,仅16.4–17.8%的轨迹保持框架一致性。不稳定的轨迹对说服性攻击的敏感度高出1.29倍(p=0.015)。在表示层面,线性探针将框架特异性编码定位到特定模型层(如Llama-3.3-70B的第63/81层,Qwen2.5-72B的第17/81层),相较训练集先验基线降低13.8–22.6%的KL散度。轻量级激活调控可减少框架整合漂移6.7–8.9%,并增强稳定性与准确性的关联。进一步提出的道德表征一致性(MRC)指标与大模型连贯性评分强相关(r=0.715,p<0.0001),其框架归因经人工标注验证,平均余弦相似度达0.859。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce \textit{moral reasoning trajectories}, sequences of ethical framework invocations across intermediate reasoning steps, and analyze their dynamics across six models and three benchmarks. We find that moral reasoning involves systematic multi-framework deliberation: 55.4--57.7\% of consecutive steps involve framework switches, and only 16.4--17.8\% of trajectories remain framework-consistent. Unstable trajectories remain 1.29$\times$ more susceptible to persuasive attacks ($p=0.015$). At the representation level, linear probes localize framework-specific encoding to model-specific layers (layer 63/81 for Llama-3.3-70B; layer 17/81 for Qwen2.5-72B), achieving 13.8--22.6\% lower KL divergence than the training-set prior baseline. Lightweight activation steering modulates framework integration patterns (6.7--8.9\% drift reduction) and amplifies the stability--accuracy relationship. We further propose a Moral Representation Consistency (MRC) metric that correlates strongly ($r=0.715$, $p<0.0001$) with LLM coherence ratings, whose underlying framework attributions are validated by human annotators (mean cosine similarity $= 0.859$).

大模型解释道德推理可解释性框架切换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。