arXiv:2607.20461cs.CLcs.AI2026-07

用情感值预测文本道德判断,为更人性化的AI伦理提供新思路。

Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study

论文配图:Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study
图 1 · 摘自论文原文
  • 构建500条人工标注的道德情感值数据集,范围-1到1。
  • 情感值与道德分类相关性显著,二分类准确率达0.764(马修相关系数)。
  • 适用于情感-道德计算研究,尤其关注人类共情与伦理决策的结合。

当前人工智能伦理系统未充分考虑情感因素。若要使AI与人类伦理对齐,需深入探究情感在人类行为、判断或陈述中的作用。尽管功利主义、德行论和康德义务论在理论上存在差异,但均不同程度关注人类情感;而描述性伦理理论中的道德基础理论则将情感置于核心地位。本文基于常识规范银行数据集,构建了一个包含500条标注的数据集,由六名参与者对行动/判断与后果的道德情感值进行评分(范围-1至1)。结果表明,情感值与多分类(不道德/酌情/道德)及二分类(不道德/道德)具有显著相关性,且通过正则化逻辑回归实现二分类的马修相关系数达0.764。这为文本道德估计提供了早期证据,表明可利用对他人影响的情感后果来推动更具人性的道德对齐AI。为促进情感-道德计算研究,本研究的标注数据将按需开放。

原文摘要 · Abstract (English)

Present implementations of artificial intelligence (AI) ethics do not adequately take feelings, or affect, into account. If AI should be aligned with human ethics, it seems reasonable to thoroughly investigate the possibility of AI behaviour that mirrors virtuous human ethical conduct, where feelings play a role in the actions, judgements or statements one makes. Furthermore, while prominent theories of normative ethics are often discussed in terms of their differences and shortcomings, Virtue, Consequentialist, and Kantian Deontological ethics all share a common feature of considering human feeling to some degree while the popular descriptive ethics theory, Moral Foundations Theory, positions feelings as central to many of its foundations. Therefore, in the present paper, a data set of moral valence is proposed, consisting of 500 annotations by six human participants for both action/judgement and consequence moral valence, ranging from -1 to 1 for text-presented scenarios from the Commonsense Norm Bank data set. The resulting valence features share significant relationships with multi-class (immoral/discretionary/moral) and binary immoral/moral categories while additionally providing a noteworthy test set Matthew's correlation coefficient of 0.764 using regularised logistic regression for binary classification. This provides early evidence of the usefulness of valence features for morality estimation of text, indicating that valenced consequences of responses for others can be considered toward more human morally-aligned AI. In the interest of promoting further affective-moral computing research, this study's annotations will be made available for research on request.

情感计算道德推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。