通过生成冲突社会规范,让AI更懂人类行为背后的道德逻辑。
Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- 用对比学习生成对立的社会规范来增强模型推理
- 在多个数据集上显著提升道德判断准确率
- 适合需要可解释伦理决策的AI系统开发者
人类行为常受社会规范约束,如‘应勇敢举报犯罪’。当前基于大规模数据训练的AI评估方法缺乏明确规范支持,难以解释且不可信。本文提出ClarityEthic,通过生成行为背后的冲突社会规范(如‘勇敢’与‘自我保护’),利用对比学习增强语言模型的道德推理能力。实验表明该方法在多个基准上优于强基线模型,人工评估也证实生成的规范能提供合理解释,有效提升伦理评估的可解释性与可信度。
原文摘要 · Abstract (English)
Human behaviors are often guided or constrained by social norms, which are defined as shared, commonsense rules. For example, underlying an action ``\textit{report a witnessed crime}" are social norms that inform our conduct, such as ``\textit{It is expected to be brave to report crimes}''. Current AI systems that assess valence (i.e., support or oppose) of human actions by leveraging large-scale data training not grounded on explicit norms may be difficult to explain, and thus untrustworthy. Emulating human assessors by considering social norms can help AI models better understand and predict valence. While multiple norms come into play, conflicting norms can create tension and directly influence human behavior. For example, when deciding whether to ``\textit{report a witnessed crime}'', one may balance \textit{bravery} against \textit{self-protection}. In this paper, we introduce \textit{ClarityEthic}, a novel ethical assessment approach, to enhance valence prediction and explanation by generating conflicting social norms behind human actions, which strengthens the moral reasoning capabilities of language models by using a contrastive learning strategy. Extensive experiments demonstrate that our method outperforms strong baseline approaches, and human evaluations confirm that the generated social norms provide plausible explanations for the assessment of human behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。