arXiv:2601.15755cs.CL2026-01ACL被引 2

评估大模型是否真实反映人口观点,不能只看回答分布,还要看观点间的关联模式。

Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs

  • 从边际分布扩展到多变量相关性,评估模型对真实人群观点结构的还原能力。
  • 真人调查数据显示,两种对齐方法均未准确复现人类观点间的实际关联模式。
  • 适合关注模型社会代表性、避免盲目乐观的研究者和评估者使用。

大型语言模型越来越多地被用于代表人类意见、价值观或信念,其向这些理想状态的可引导性是当前研究热点。现有工作主要聚焦于对齐边际响应分布,将每个评估样本独立处理。尽管必要,但可能忽略表征真实人群的深层潜在结构,而这些结构支撑着文化价值观理论。本文提出一个框架,通过多变量相关性模式而非仅边际分布来评估对齐模型的代表性。我们通过对比两种模型引导技术(人物提示与人口细分微调)在世界价值观调查(World Values Survey)人类响应上的表现,验证该评估方案的价值。结果显示:人口细分微调模型更接近边际响应分布,而人物提示在复现调查题项间的实证相关结构上略优。然而,两种方法均未真正匹配人类的相关模式。结论表明,代表性是价值对齐的一个独立维度,仅关注边际分布可能掩盖结构性缺陷,导致对模型代表性的过度乐观判断。

原文摘要 · Abstract (English)

Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an active area of research. Existing work focuses predominantly on aligning marginal response distributions, treating each alignment evaluation example independently. While essential, this may overlook deeper latent structures that characterise real populations and underpin cultural values theories. We propose a framework for evaluating the \textit{representativeness} of aligned models through multivariate correlation patterns in addition to marginal distributions. We show the value of our evaluation scheme by comparing two model steering techniques (persona prompting and demographic fine-tuning) and evaluating them against human responses from the World Values Survey. While the demographic fine-tuned model better approximates marginal response distributions, persona prompting performs marginally better at reproducing the empirical correlation structure between survey items. Despite this reversal, neither technique aligns with human correlation patterns. We conclude that representativeness is a distinct aspect of value alignment and an evaluation focused on marginals can mask structural failures, leading to overly optimistic conclusions about model representativeness.

大模型评估价值观对齐代表性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。