arXiv:2607.05113cs.CLcs.CY2026-07中稿 · EMNLP

用户对大模型的评价,更多反映预期而非实际表现。

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

论文配图:Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
图 1 · 摘自论文原文
  • 通过前后暗示改变用户对模型能力的预期,影响评分与交互方式。
  • 实际任务表现不变时,夸大预期者评分更高、指令更直接。
  • 评价结果主要衡量期望满足度,而非模型真实性能,适合评估设计者。

想象两位用户使用同一款大模型:一人被告知是前沿旗舰,另一人则认为是旧版弱模型。尽管使用相同模型,两人对模型有用性和智能性的评分却差异显著。在一项受控研究中,162名参与者在三个协作任务前,先浏览了匹配、夸大或低估其模型真实能力的落地页。这种预交互框架改变了用户印象与行为,但任务表现未变。被夸大的用户评分更积极,指令更直接;被低估的用户则撰写更长、更协作的提示。用户与模型共同产出的质量仅取决于模型真实能力,不受预期影响。用户使用后印象变化(两种测量方式)不被任务表现预测(β = -0.01 与 0.11,均不显著),却显著由是否满足预期(β = 0.47 与 0.50,p < .001)和操作信心(β = 0.47 与 0.36,p < .001)决定。交互后,用户仍在评价‘宣传’而非‘产品’:当前用户生成的模型评价,包括驱动公开排行榜的偏好数据,至少与模型本身一样,反映的是期望管理。

原文摘要 · Abstract (English)

Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 participants each used one of six LLMs from two families across three collaborative tasks, after first viewing a landing page that matched, overstated, or understated their model's true capability. This pre-interaction framing shifted user opinions and interaction behavior while task performance did not. Oversold users rated the model more favorably and used more directive prompting, while Undersold users wrote longer, more collaborative prompts. The quality of what users and the model produced together depended only on the model's true capability, not on what users were told. Participants' change in model impressions after use, measured across two impression measures, was not predicted by task performance ($β= -0.01$ and $0.11$, both n.s.), but by whether the model met users' expectations ($β= 0.47$ and $0.50$, both $p < .001$) and how confident they felt working with it ($β= 0.47$ and $0.36$, both $p < .001$). After interaction, users are still rating the pitch, not the product: user-elicited LLM evaluations, including the preference data driving public leaderboards, measure expectation management at least as much as the model itself.

大模型评估用户心理期望效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。