揭示大模型道德推理的困境:现有训练方法难突破语用瓶颈。
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
- 基于语用学分析,发现道德推理受话语隐含语境制约。
- 当前训练范式下,模型泛化能力受限,难以跨场景迁移。
- 适合关注大模型价值观对齐与伦理安全的研究者阅读。
确保大语言模型(LLMs)输出符合社会价值的内容,对其广泛应用至关重要。以往研究表明,LLMs在需要道德认知的任务(如伦理判断)上表现不佳。尽管现有方法通过微调模型以提升其能力,但如何选择最优学习范式仍存在争议。本文从分布语义理论和道德话语的语用特性出发,分析发现性能提升机制与语义任务相似,受话语中隐含的道德语用特征影响,这一现象称为‘语用困境’。结论表明,该语用困境严重限制了现有学习范式的泛化能力,是当前大模型获取道德推理能力的主要瓶颈。
原文摘要 · Abstract (English)
Ensuring that Large Language Models (LLMs) return just responses which adhere to societal values is crucial for their broader application. Prior research has shown that LLMs often fail to perform satisfactorily on tasks requiring moral cognizance, such as ethics-based judgments. While current approaches have focused on fine-tuning LLMs with curated datasets to improve their capabilities on such tasks, choosing the optimal learning paradigm to enhance the ethical responses of LLMs remains an open research debate. In this work, we aim to address this fundamental question: can current learning paradigms enable LLMs to acquire sufficient moral reasoning capabilities? Drawing from distributional semantics theory and the pragmatic nature of moral discourse, our analysis indicates that performance improvements follow a mechanism similar to that of semantic-level tasks, and therefore remain affected by the pragmatic nature of morals latent in discourse, a phenomenon we name the pragmatic dilemma. We conclude that this pragmatic dilemma imposes significant limitations on the generalization ability of current learning paradigms, making it the primary bottleneck for moral reasoning acquisition in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。