测试大模型对新闻标题情感倾向的理解能力,发现不同人群反应差异大。
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

- 用英国3011人问卷和7个大模型对比情感判断
- 顶尖模型与人类相关性达0.789,但不同人群差异显著
- 提醒开发者关注不同群体的适配性,避免偏见
大型语言模型(LLMs)正深刻影响人们获取信息和形成世界观的方式。除了已知的偏见问题,当前更需关注的是:这些模型能否理解文本框架所传递的情感细微差别?本研究通过实证方法评估多个大模型在情感感知上的对齐程度。基于覆盖政治与地缘政治冲突的新闻标题,采用英国代表性成人样本(n=3011,来自YouGov调查)与七个大模型共同判断标题是否引发对特定一方的同情。结果显示,各模型与人类评价的相关性从0.789(GPT-5.2)到0.4(Mistral Large 2512)不等。关键在于,尽管领先模型整体上与人类判断一致,涵盖年龄、性别、教育水平、地缘政治知识及立场偏好等人口学子群组,但仍存在统计显著差异。该研究以严谨设计和大规模多样化数据集,提供了迄今最全面的大模型新闻框架理解能力评估。结果揭示了一个常被忽视的重要问题:即使总体对齐度高,模型表现仍可能因人群特征和文化规范而异。因此,在开发伦理化、实用型AI系统时,是否考虑这种差异化对齐,将产生重大影响。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。