用概率分布建模多元伦理,让AI决策更接近人类道德判断。
Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

- 构建伦理单纯形框架,融合三大主流伦理理论及15个子理论。
- 在450个案例上实现88.89%准确率,显著优于传统二元判断。
- 适合研究AI伦理对齐、可解释性与道德分歧分析的学者。
社会关键领域中人工智能决策日益普及,但现有方法多采用标量或二元道德判断,缺乏解释力且忽视情境与理论背景。为此,我们提出一种将道德推理建模为规范伦理理论分布的方法,引入规范伦理单纯形以整合多种理论。构建了包含450个案例、覆盖15个细粒度子理论的基准数据集,案例以自然语言描述伦理困境并附有提取的情境特征。通过双流规范-语义架构实现单纯形建模,并采用序列堆叠集成学习,优化对后果主义、德行伦理与义务论三大理论及其子类的匹配。实验表明,结合情境与规范先验及语义嵌入后,分类准确率达88.89%。消融实验证明结构化伦理表示超越类比推理,堆叠架构因逐步学习粒度而表现最优。通过熵、置信度与可视化分析伦理多元性。该方法支持类人道德推理、道德分歧解析与未来AI对齐。
原文摘要 · Abstract (English)
Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of autonomous systems, most approaches to handling autonomous moral decision-making resort to scalar or binary judgments. These methods are insufficient for acceptable moral reasoning, as they provide little explanation, leaving out imperative contextual and theoretical information that must be included to support accountability. For this, we propose a framework to model moral reasoning as a distribution over normative ethical theories or ethical pluralism. We introduce a normative ethics simplex that integrates these theories. A benchmark of 450 cases across 15 fine-grained subtheories was also prepared for the purposes of stacked ensemble learning. These cases describe ethical dilemmas in natural language and have associated extracted contextual features. The implementation of the simplex was achieved via a two-stream normative-semantic architecture. This is followed by the fusion of normative information and a sequential, stacking ensemble to learn the best fit of the three broad theories: consequentialism, virtue ethics, and deontology, and the 15 subcategories. Our experiments demonstrate that the integration of contextual and normative priors with the semantic embeddings significantly improves the performance of the classification, displaying an accuracy of 88.89%. We conducted ablation studies to show that structured ethical representations contribute beyond analogical reasoning, and the chosen stacking architecture gives the best results due to the gradual learning of granularity. Ethical pluralism is also analyzed through entropy, confidence, and visualization. Thus, modeling ethical pluralism as a probabilistic normative distribution supports human-like moral reasoning, ethical disagreement analysis, and future alignment in AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。