研究发现语言模型的表态与行为偏好不一致,关键在于提问方式。
Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models
- 通过允许中立和放弃回答,提升表态与行为偏好的一致性
- 在行为选择中允许放弃时,一致性反而大幅下降
- 用表态结果引导行为选择无法稳定提升一致性
近期研究发现语言模型存在表态-行为偏好差距:模型宣称的价值观与其实际选择不一致。现有评估多依赖二元强制选择提示,将真实偏好与提示机制的干扰混为一谈。本文系统考察了24个语言模型在不同提问协议下的偏好一致性。在表态阶段允许中立或放弃回答,可有效排除弱信号,显著提升表态与强制选择行为之间的斯皮尔曼等级相关系数(ρ)。然而,若在行为选择中也允许放弃,ρ值会降至接近零甚至为负,因中立率过高。此外,在行为评估中使用表态结果作为系统提示引导,对AIRiskDilemmas数据集上的一致性无可靠提升。结果表明,偏好一致性高度依赖于提问协议,偏好获取需考虑不确定偏好情况。
原文摘要 · Abstract (English)
Recent work identifies a stated-revealed (SvR) preference gap in language models (LMs): a mismatch between the values models endorse and the choices they make in context. Existing evaluations rely heavily on binary forced-choice prompting, which entangles genuine preferences with artifacts of the elicitation protocol. We systematically study how elicitation protocols affect SvR correlation across 24 LMs. Allowing neutrality and abstention during stated preference elicitation allows us to exclude weak signals, substantially improving Spearman's rank correlation ($ρ$) between volunteered stated preferences and forced-choice revealed preferences. However, further allowing abstention in revealed preferences drives $ρ$ to near-zero or negative values due to high neutrality rates. Finally, we find that system prompt steering using stated preferences during revealed preference elicitation does not reliably improve SvR correlation on AIRiskDilemmas. Together, our results show that SvR correlation is highly protocol-dependent and that preference elicitation requires methods that account for indeterminate preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。