用语用推理提升大模型道德判断泛化能力
From Training to Generalization: Improving Moral Reasoning Through Pragmatic Inference
- 结合语用链接与道德基础理论,让模型理解情境中的道德规范
- 在未见测试数据上显著提升道德推理泛化性能
- 适合关注伦理智能、可解释AI的研究者
尽管道德推理已成为大语言模型(LLMs)的前沿研究方向,但其泛化能力仍存挑战:模型在训练数据上表现良好,却难以推广到未见测试数据。从语言学角度看,道德推理是一种语用过程,需基于社会规范语境推断道德判断。然而现有方法忽视了这一语用特性,主要受制于两大瓶颈:(1) LLMs擅长捕捉分布语义,而道德推理依赖语用理解;(2) 缺乏有效手段将语言锚定于道德语境。本文提出一种语用推理方法,通过结合元语用链接与道德基础理论(Moral Foundations Theory),使模型在给定道德情境下推断道德判断。其中,元语用链接弥合分布语义与语用之间的鸿沟,而道德基础理论为语言提供道德语境的规范依据。实验表明,该方法显著提升了大模型在道德推理上的泛化能力,凸显语用推理在未来研究中的潜力。
原文摘要 · Abstract (English)
Although moral reasoning has emerged as a promising research direction for large language models (LLMs), a persistent generalization challenge remains: LLMs often achieve strong performance on training data but struggle to generalize their moral reasoning to unseen test data. From a linguistic perspective, moral reasoning is a pragmatic process in which moral judgments are inferred based on the context of social norms underlying a given moral situation. However, existing approaches overlook this pragmatic nature because of two major bottlenecks: (1) LLMs are primarily skilled in capturing distributional semantics, which differs from the pragmatic nature of moral reasoning; (2) there is currently no effective solution for grounding language in the moral context. In this paper, we develop a pragmatic inference approach that enables LLMs to infer moral judgments for a given moral situation by combining metapragmatic links with Moral Foundations Theory. Specifically, metapragmatic links serve to bridge the gap between distributional semantics and pragmatics, whereas Moral Foundations Theory provides a principled basis for grounding language in moral contexts. Experimental results demonstrate that our approach substantially improves LLMs' generalization in moral reasoning, highlighting the potential of pragmatic inference for future moral reasoning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。