arXiv:2605.03609cs.AIcs.LG2026-05ACL

通过精准干预模型内部道德路径,实现可解释的伦理偏好调控。

Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models

论文配图:Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 定位变压器中道德决策的交汇与分流点,选择性关闭非目标路径。
  • 在真实道德困境测试中,精准校准伦理偏好且保持模型通用能力。
  • 适合需要可控伦理输出的AI应用开发者或研究者使用。

大型语言模型在不同情境下常表现出不一致的道德偏好。本文研究推理时对特定伦理框架的引导,同时保持模型整体能力。提出「收敛-发散路由」机制,识别并编辑变换器块中道德相关路径首次交汇后又分叉的关键节点。在这些位置屏蔽非目标分支,阻止下游传播但保留上游计算。实验表明该干预显著提升目标伦理框架的推理表现。为实现细粒度控制,将共空间模式方法应用于残差流,提取每层分支点处区分功利主义与义务论框架的两个方向。进而引入双逻辑校准,一种闭式、最小ℓ₂范数更新,将残差投影至该二维子空间,使方向投影与用户指定偏好权重对齐。在真实道德困境数据集上的实验显示,本方法能可靠实现偏好校准,显著优于近期基线,且保持良好通用性,提供可解释的调控机制。

原文摘要 · Abstract (English)

Large language models often display heterogeneous moral preferences across settings. We study inference-time steering toward a desired ethical framework while preserving general competence. We present Convergent-Divergent Routing, which traces and edits minimal branch points inside transformer blocks where ethical-framework-related pathways first converge and then diverge. Gating non-target branches at these loci blocks the downstream propagation while leaving upstream computations intact. We find that this intervention alone increases targeted ethical-framework reasoning. To achieve fine-grained control, we adapt Common Spatial Patterns to the residual stream and extract, for each branch-point layer, a pair of directions that discriminate between utilitarian and deontological frameworks. We then introduce Dual Logit Calibration, a closed-form, minimum-$\ell_2$-norm update that moves the residual within this two-dimensional subspace so the resulting directional projections align with user-specified preference weights. Experiments on real-life moral dilemmas show that our method reliably achieves preference calibration and largely preserves general capabilities, outperforming recent baselines while providing an interpretable mechanism.

伦理对齐可控生成模型解释大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。