用认知忠实模型提升AI决策对齐效果
Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- 基于成对比较学习人类决策的结构化认知过程
- 在肾源分配任务中准确率超越已有模型
- 适合关注可解释性与人类对齐的研究者
当前人工智能趋势致力于将模型对齐于人类中心目标,如个人偏好、效用或社会价值。传统偏好获取方法常无法捕捉人类决策背后的认知机制,如启发式或简化思维模式。为此,本文采用公理化方法,从成对比较中学习认知忠实的决策过程。基于认知决策研究,我们构建了一类模型:先通过学习规则处理特征,再以固定规则(如Bradley-Terry规则)聚合生成决策。这种结构化信息处理使模型更贴近真实人类决策。我们在肾源分配任务中训练出可解释模型,结果表明其准确率与现有模型相当或更优。
原文摘要 · Abstract (English)
Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using standard preference elicitation methods, researchers and practitioners build models of human decisions and judgments, to which AI models are aligned. However, standard elicitation methods often fail to capture the cognitive processes behind human decision making, such as heuristics or simplifying structured thought patterns. To address this failure, we take an axiomatic approach to learning cognitively faithful decision processes from pairwise comparisons. Building on the literature analyzing cognitive processes that shape human decision-making, we derive a model class in which features are first processed with learned rules, then aggregated via a fixed rule, such as the Bradley-Terry rule, to produce a decision. This structured processing of information ensures that such models are realistic and feasible candidates to represent underlying human decision-making processes. We demonstrate the efficacy of this modeling approach by learning interpretable models of human decision making in a kidney allocation task, and show that our proposed models match or surpass the accuracy of prior models of human pairwise decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。