arXiv:2605.08556cs.LG2026-05被引 2

通过用户行为反推大模型的真实偏好,检验其对齐与可引导性。

Can Revealed Preferences Clarify LLM Alignment and Steering?

  • 基于模型决策和概率输出,拟合其隐含成本函数。
  • 多模型在医疗诊断中展现部分一致性,但对指令响应不稳。
  • 适合关注大模型价值观对齐与可控性的研究者。

大语言模型在不确定性情境下越来越多地参与或支持高风险决策,其对齐不仅依赖事实准确性,还取决于模型如何权衡不同结果。我们提出一个实证流程,用于估计模型在观察到的选择中所优化的隐含偏好:通过获取模型对未知项的概率分布及其在决策任务中的选择,再拟合离散选择模型以恢复最能解释模型决策的成本函数。该方法使我们能够严格评估模型是否以一致的目标导向方式行动、能否用语言描述与其决策策略匹配的目标,以及提示是否能可靠引导其政策以实现用户指定的成本函数。我们在四个医疗诊断领域及多个前沿与开源模型上应用此评估。结果显示,尽管许多模型具有非平凡的内部一致性,但在忠实回应或采纳用户方向的偏好方面仍存在显著缺陷。

原文摘要 · Abstract (English)

LLMs are increasingly used to make or support high-stakes decisions under uncertainty, where alignment depends not only on factual accuracy but on how models weigh tradeoffs between different outcomes. We present an empirical pipeline for estimating the implied preferences that an LLM's observed choices optimize: we elicit the model's probability distribution over unknowns along with the choice it would make for the decision task and then fit a discrete choice model to recover the cost function that best rationalizes the model's decisions. We show how this revealed-preference description allows rigorous evaluation of whether models behave in a consistently goal-directed way, whether they can verbalize a description of their objectives which matches their revealed decision policy, and whether prompting can reliably steer those policies to implement a user-specified cost function. We apply this evaluation across four medical diagnosis domains and multiple frontier and open-source models. We find that while many models have a nontrivial degree of internal coherence, they also have significant weaknesses in faithfully reporting or adopting preferences in response to user direction.

模型对齐偏好建模医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。