LLM给出的偏好判断不自洽,无法用单一效用函数描述。
LLM-Derived Preference Judgments Are Not Self-Consistent

- 通过统计检验衡量LLM偏好判断的自洽性
- 六种LLM在航班、公寓等场景中均存在显著不一致
- 适合研究人机交互与决策建模的学者参考
智能体越来越多地通过查询大语言模型(LLM)获取数值型偏好判断来理解人类自然语言偏好,例如询问某人愿意为某物品支付多少金额。现有方法基于这些判断估计效用函数,并据此选择最优行动。该流程隐含假设:偏好判断具有近似自洽性——即存在一个单一效用函数可复现所有判断。但这一假设是否成立?我们测量了基数型LLM偏好判断的自洽性。例如,两件物品间声称的支付意愿差值,应等于使个体对交换无差异的支付金额。我们开发了统计检验与可解释的偏离度量,用于评估实际响应与最佳拟合自洽效用函数之间的差距。在六种不同LLM上,针对航班、公寓和酒店等示例进行实验,发现不一致现象普遍存在且显著。这表明LLM生成的偏好判断无法被单一效用函数忠实概括。
原文摘要 · Abstract (English)
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。