用人类道德判断数据构建可解释的道德情境模型,提升对模糊行为评价的准确性。
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
- 通过概率聚类与大模型语义抽象结合,从人类判断中学习行为的情境特征。
- 在60%的案例中比直接使用大模型更贴近多数人判断,显著提升预测准确率。
- 适合需要透明决策依据的AI伦理、自动驾驶等需解释性判断的场景。
道德行为的评判不仅取决于结果,还依赖于具体情境。我们提出COMETH(基于文本人类输入的道德评价情境组织框架),融合概率情境学习、大语言模型语义抽象与人类道德判断,建模情境如何影响模糊行为的可接受性。我们构建了一个基于实证的数据集,包含300个情境,涵盖六种核心行为(如‘不杀人’、‘不欺骗’、‘不违法’),并收集了101名参与者对每种行为的三元判断(责备/中立/支持)。通过大模型过滤和MiniLM嵌入配合K-means聚类,标准化动作并生成稳定的核心行为簇。COMETH在线聚类人类判断分布,利用合理的差异准则学习特定行为的道德情境。为实现泛化与解释,其泛化模块提取简洁的非评价性二元情境特征,并在透明的概率模型中学习特征权重。实验表明,相较于端到端的大模型提示,COMETH在平均上将与多数人判断的一致性提高约一倍(约60%对比约30%),同时揭示影响预测的关键情境因素。贡献包括:(i) 一个基于实证的道德情境数据集;(ii) 一套结合人类判断与模型情境学习的可复现流程;(iii) 一种可解释的替代方案,用于情境敏感的道德预测与解释。
原文摘要 · Abstract (English)
Moral actions are judged not only by their outcomes but by the context in which they occur. We present COMETH (Contextual Organization of Moral Evaluation from Textual Human inputs), a framework that integrates a probabilistic context learner with LLM-based semantic abstraction and human moral evaluations to model how context shapes the acceptability of ambiguous actions. We curate an empirically grounded dataset of 300 scenarios across six core actions (violating Do not kill, Do not deceive, and Do not break the law) and collect ternary judgments (Blame/Neutral/Support) from N=101 participants. A preprocessing pipeline standardizes actions via an LLM filter and MiniLM embeddings with K-means, producing robust, reproducible core-action clusters. COMETH then learns action-specific moral contexts by clustering scenarios online from human judgment distributions using principled divergence criteria. To generalize and explain predictions, a Generalization module extracts concise, non-evaluative binary contextual features and learns feature weights in a transparent likelihood-based model. Empirically, COMETH roughly doubles alignment with majority human judgments relative to end-to-end LLM prompting (approx. 60% vs. approx. 30% on average), while revealing which contextual features drive its predictions. The contributions are: (i) an empirically grounded moral-context dataset, (ii) a reproducible pipeline combining human judgments with model-based context learning and LLM semantics, and (iii) an interpretable alternative to end-to-end LLMs for context-sensitive moral prediction and explanation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。