让大模型学会诊断并纠正道德错误,而非仅模仿正确表述。
Learning to Diagnose and Correct Moral Errors: Beyond Shallow Heuristics in Moral Alignment
- 通过语用推理机制,让模型理解道德判断背后的逻辑
- 在跨任务场景中显著提升道德对齐效果,泛化能力更强
- 适合研究伦理对齐、AI安全的学者与工程师
现有道德对齐方法主要将大语言模型(LLM)生成内容与合乎道德的语言分布对齐,虽取得一定进展,但往往脆弱且依赖浅层启发式规则,在分布外任务上表现下降。本质上,这些方法只教会模型‘什么是’道德上合适或不合适的话语,而非‘为什么’如此。本文提出基于语用推理的诊断与修正方法,使模型能够理解道德话语背后的隐含意义。该方法根据不同道德论述的推理负载调整推断过程,而非分别建模其复杂的语义分布。实验证明,本方法显著提升了模型的道德对齐能力,并在多种任务间具有良好泛化性能。
原文摘要 · Abstract (English)
Existing approaches to moral value alignment are primarily set out to align LLMs' generation with the distributions of morally appropriate language, which has seen good progress. However, these approaches are often brittle, heavily rely on shallow heuristics, and reduce performance in out-of-the-distribution tasks. In other words, the learning paradigm underlying existing approaches teaches LLMs what morally (in)appropriate language looks like, but not why it is morally (in)appropriate. In this paper, we address this challenge by developing pragmatic inference-driven methods to facilitate LLMs' learning of how to diagnose and correct moral errors, thereby enabling them to generate morally appropriate language. Pragmatic inference is the reasoning process of deriving (implied) meanings -- a famous concept in linguistics. Our methods vary the inference procedures by the inferential load of different moral discourses, rather than modelling their diverse and complex semantic distributions separately. Empirical results demonstrate that our approach improves moral value alignment in LLMs and generalizes effectively across tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。