用道德理论解析大模型内在道德观,实现精准安全干预。
The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
- 基于道德基础理论构建可操控的道德向量
- 跨语言验证道德表征共享性,提升泛化能力
- 动态融合检测与注入,平衡安全与帮助性
提升大语言模型(LLMs)的道德对齐是人工智能安全的关键挑战。现有对齐方法常仅作为表面防护,未触及模型内在道德表征。本文借助道德基础理论(MFT),映射并操纵大模型细粒度的道德空间。通过跨语言线性探测,验证了中层语义中道德表征的共性,并发现英汉语言间存在共享但不同的道德子空间。在此基础上,提取可调节的道德向量,并在内部与行为层面均验证其有效性。利用道德的高泛化性,提出自适应道德融合(AMF)——一种推理时动态干预机制,结合探测与向量注入,解决安全与有用性之间的权衡。实验表明,该方法作为靶向内在防御,显著降低良性查询的错误拒绝率,同时有效抑制越狱攻击成功率,优于标准基线。
原文摘要 · Abstract (English)
Enhancing the moral alignment of Large Language Models (LLMs) is a critical challenge in AI safety. Current alignment techniques often act as superficial guardrails, leaving the intrinsic moral representations of LLMs largely untouched. In this paper, we bridge this gap by leveraging Moral Foundations Theory (MFT) to map and manipulate the fine-grained moral landscape of LLMs. Through cross-lingual linear probing, we validate the shared nature of moral representations in middle layers and uncover a shared yet different moral subspace between English and Chinese. Building upon this, we extract steerable Moral Vectors and successfully validate their efficacy at both internal and behavioral levels. Leveraging the high generalizability of morality, we propose Adaptive Moral Fusion (AMF), a dynamic inference-time intervention that synergizes probe detection with vector injection to tackle the safety-helpfulness trade-off. Empirical results confirm that our approach acts as a targeted intrinsic defense, effectively reducing incorrect refusals on benign queries while minimizing jailbreak success rates compared to standard baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。