用对比学习让大模型识别情境中的社会规范,提升道德判断可信度。
ClarityEthic: Explainable Moral Judgment Utilizing Contrastive Ethical Insights from Large Language Models
- 通过对比不同视角的社会规范,筛选最可靠的判断依据。
- 在多个道德判断任务中优于现有方法,人类评估认可解释合理性。
- 适合研究大模型伦理对齐与可解释决策的学者使用。
随着大型语言模型(LLMs)的广泛应用,确保其安全性以避免对人类造成伤害并促进伦理行为至关重要。然而,直接依赖大规模数据训练来判断价值取向(支持或反对)既不可靠也不可解释。我们假设,模仿人类依靠社会规范做道德决策,有助于大模型理解并预测道德判断。但捕捉人类价值观仍具挑战,因多种相关规范在特定情境下可能冲突。例如,多数人认同且有利于社会福祉的规范(如“不作弊”)更易被接受和采纳。因此,大模型在做出道德判断前,需识别出适用的具体规范。为此,我们提出一种名为ClarityEthic的新方法,利用大模型的推理能力与对比学习,从多角度挖掘人类行为相关的社会规范,并选择最可靠的规范以提升判断准确率。大量实验表明,该方法在道德判断任务中优于当前最优方法。此外,人工评估证实生成的社会规范提供了合理且可信的解释。这表明,通过模拟人类道德策略建模大模型的道德判断,是提升其伦理行为的有效路径。
原文摘要 · Abstract (English)
With the rise and widespread use of Large Language Models (LLMs), ensuring their safety is crucial to prevent harm to humans and promote ethical behaviors. However, directly assessing value valence (i.e., support or oppose) by leveraging large-scale data training is untrustworthy and inexplainable. We assume that emulating humans to rely on social norms to make moral decisions can help LLMs understand and predict moral judgment. However, capturing human values remains a challenge, as multiple related norms might conflict in specific contexts. Consider norms that are upheld by the majority and promote the well-being of society are more likely to be accepted and widely adopted (e.g., "don't cheat,"). Therefore, it is essential for LLM to identify the appropriate norms for a given scenario before making moral decisions. To this end, we introduce a novel moral judgment approach called \textit{ClarityEthic} that leverages LLMs' reasoning ability and contrastive learning to uncover relevant social norms for human actions from different perspectives and select the most reliable one to enhance judgment accuracy. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches in moral judgment tasks. Moreover, human evaluations confirm that the generated social norms provide plausible explanations that support the judgments. This suggests that modeling human moral judgment with the emulating humans moral strategy is promising for improving the ethical behaviors of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。