arXiv:2511.20689q-bio.NCcs.AI2025-11被引 2

将道德认知嵌入大模型架构,从底层设计提升AI的伦理决策能力。

Morality in AI. A plea to embed morality in LLM architectures and frameworks

  • 以动态注意力机制为核心,重构模型结构与处理间的接口。
  • 基于穆尔多克的'爱之关注'理论,提出三类可操作的道德嵌入路径。
  • 适合关注AI伦理、架构设计及跨学科合作的研究者参考。

大型语言模型(LLMs)日益参与人类决策与行为引导,确保其对道德意义的处理已成为关键挑战。当前方法主要依赖微调和人类反馈强化学习等自下而上的策略。本文提出一种根本性新思路:通过自上而下的设计原则,将道德意义处理直接嵌入基于Transformer的模型架构与框架中。我们首先构建一个框架,将注意力视为连接结构与处理的动态界面,区别于心理学中的线性注意力模型。借鉴神经架构设计中的生物-人工注意力类比,优化认知处理能力。进一步将此分析延伸至道德处理,运用伊里斯·穆尔多克的'爱之关注'理论(持续、公正的观察,能以清晰与共情重塑对他人认知)来探讨人类与大模型道德处理的功能类比。我们提出了若干可能有效的技术实现方式,并进行评估。承认探索局限性的同时,提出三项核心贡献:(1) 将注意力概念化为连接结构与处理的动态系统机制;(2) 基于穆尔多克的‘爱之关注’,提出通过修改训练目标、运行时权重调整及注意力架构优化实现道德嵌入的技术路径;(3) 主张将道德融入架构与框架,可补充外部约束型方法。最后呼吁变压器架构设计师与从事人工智能伦理的哲学家开展协作。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mediate human decision-making and behaviour. Ensuring LLM processing of moral meaning therefore has become a critical challenge. Current approaches rely predominantly on bottom-up methods such as fine-tuning and reinforcement learning from human feedback. We propose a fundamentally different approach: embedding moral meaning processing directly into the architectural mechanisms and frameworks of transformer-based models through top-down design principles. We first sketch a framework that conceptualizes attention as a dynamic interface mediating between structure and processing, contrasting with existing linear attention frameworks in psychology. We start from established biological-artificial attention analogies in neural architecture design to improve cognitive processing. We extend this analysis to moral processing, using Iris Murdoch's theory of loving attention (sustained, just observation that enables moral transformation by reseeing others with clarity and compassion) to philosophically discuss functional analogies between human and LLM moral processing. We formulate and evaluate potentially promising technical operationalizations to embed morality in LLM architectures and frameworks. We acknowledge the limitations of our exploration and give three key contributions. (1) We conceptualize attention as a dynamic system mechanism mediating between structure and processing. (2) Drawing on the Murdoch notion of loving attention, we outline technical pathways for embedding morality in LLMs, through modified training objectives, runtime weight adjustments, and architectural refinements to attention. (3) We argue that integrating morality into architectures and frameworks complements external, constraint-based methods. We conclude with a call for collaboration between transformer designers and philosophers engaged in AI ethics.

AI伦理注意力机制架构设计道德认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。