用大模型生成道德反馈,让智能体在复杂环境中做出更可靠的伦理决策。
Addressing Moral Uncertainty using Large Language Models for Ethical Decision-Making
- 用大模型模拟多种伦理理论,自动给出行为评分
- 通过信念融合算法整合多视角道德判断,生成可训练的奖励信号
- 减少人工设计奖励,适合应对突发伦理挑战的现实场景
我们提出一种伦理决策框架,通过任务无关的道德层对预训练强化学习(RL)模型进行优化。初始训练后,采用由大语言模型(LLM)生成的反馈替代人工反馈进行伦理微调。该LLM融合后果主义、义务论、美德论、社会正义与关怀伦理等多重道德原则,为推荐行为赋予信念值。道德层利用信念詹森-香农散度与德普斯特-沙弗理论,聚合多个来自LLM的道德视角得分,生成概率分数作为塑造奖励,引导智能体选择符合平衡伦理框架的行为。该集成学习框架帮助RL智能体在复杂环境中应对道德不确定性,实现跨任务的道德决策。实验对比不同LLM变体及其它信念聚合方法,结果表明本方法在一致性、适应性上均有提升,且显著降低对手工设计伦理奖励的依赖。该方法尤其适用于突发伦理挑战的动态场景,具备良好的真实应用潜力。
原文摘要 · Abstract (English)
We present an ethical decision-making framework that refines a pre-trained reinforcement learning (RL) model using a task-agnostic ethical layer. Following initial training, the RL model undergoes ethical fine-tuning, where human feedback is replaced by feedback generated from a large language model (LLM). The LLM embodies consequentialist, deontological, virtue, social justice, and care ethics as moral principles to assign belief values to recommended actions during ethical decision-making. An ethical layer aggregates belief scores from multiple LLM-derived moral perspectives using Belief Jensen-Shannon Divergence and Dempster-Shafer Theory into probability scores that also serve as the shaping reward, steering the agent toward choices that align with a balanced ethical framework. This integrated learning framework helps the RL agent navigate moral uncertainty in complex environments and enables it to make morally sound decisions across diverse tasks. Our approach, tested across different LLM variants and compared with other belief aggregation techniques, demonstrates improved consistency, adaptability, and reduced reliance on handcrafted ethical rewards. This method is especially effective in dynamic scenarios where ethical challenges arise unexpectedly, making it well-suited for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。