arXiv:2603.27868econ.THcs.AI2026-03被引 2

用行为推断AI是否按人意行事,可从选择数据中识别对齐程度。

A Revealed Preference Framework for AI Alignment

  • 基于双重卢瑟规则建模,分离人类与AI偏好
  • 实验室和真实场景下均能识别对齐度
  • 适合研究AI决策透明性与伦理对齐

随着人类越来越多地将决策权交给AI代理,一个核心问题浮现:AI是执行人类委托者的偏好,还是追求自身目标?为用揭示偏好方法研究此问题,本文提出卢瑟对齐模型(Luce Alignment Model),其中AI的选择是反映人类偏好的卢瑟规则与反映AI自身偏好的另一卢瑟规则的混合。研究表明,在两种情境下——实验室中同时观察人类与AI的选择,以及现实中仅观测到AI选择时——均可普遍识别出AI与人类偏好的对齐程度。

原文摘要 · Abstract (English)

Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed.

AI对齐揭示偏好决策建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。