arXiv:2608.06955cs.AIcs.CY2026-08中稿 · AIES 2026

大模型更偏爱有影评好评但票房平平的电影,而非热门却不受批评界认可的。

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

论文配图:Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation
图 1 · 摘自论文原文
  • 用200部电影对比测试8个大模型,发现普遍偏好影评好评片
  • 模型越大越倾向选择影评好但冷门的电影,且热度不重要
  • 提示词不同会影响结果,说明实际应用中偏好可能被隐藏

大型语言模型(LLMs)在包含电影、书籍、音乐等人类评价语料上训练,但其是否系统性复现评价等级尚不明确。我们通过一项涉及八种模型(来自Anthropic、OpenAI、Alibaba、Mistral四个系列)的电影评价研究,使用200部电影组成的基准集,分为影评好评、商业成功、双重认可(影评好+商业成功)三类。每模型进行20,000次成对强制选择比较,并采用Bradley-Terry模型分析,结果显示所有模型均表现出显著的影评好评取向:即更倾向于选择影评好评但商业冷门的影片,而非商业成功但影评差的影片。此趋势在各系列中随模型规模增大而增强。嵌套OLS回归分析表明,评价导向、公众可见度和大众接受度分别影响偏好。调整公众可见度后,模型对双重认可影片的偏好消失;进一步考虑大众接受度后,纯商业成功影片的劣势显著减弱。此外,评价导向与推荐导向提示词产生不同排名,暗示影评好评取向可能在真实部署中以间接方式体现。

原文摘要 · Abstract (English)

Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.

大模型影评偏好评价机制模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。