用精确的统计方法测量大模型对国家的偏见,突破传统评估局限。
Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

- 设计交叉实验,分离出模型偏见的主效应与交互作用。
- 直接操作令牌级概率分布,消除采样噪声,结果更精准。
- 适用于研究模型系统性偏见,尤其适合做公平性分析的人。
随着大型语言模型(LLMs)越来越多地作为自主代理部署,准确评估其潜在价值观和偏见至关重要。当前自然语言处理领域通常使用大规模、非结构化的基准测试来评估模型性能。尽管这些数据集在评估通用能力方面有效,但它们本质上混淆了因果机制:即使检测到整体偏见,非结构化评估也无法区分该偏见是源于基础特质、上下文混杂因素,还是复杂交互作用。为解决此问题,我们提出一种分析上精确的框架,用于对LLM进行受控行为评估。通过将人类心理测量学与LLM机制相结合,填补了设计、测量与分析方面的空白。首先,我们用完全交叉的因子实验替代非结构化提示,系统性地分离因果主效应与交互效应。其次,通过直接操作精确的令牌级概率质量函数(PMFs),消除蒙特卡洛文本采样带来的噪声。第三,我们推导出一种多变量序数共识度量与分布型方差分析(ANOVA),以解析这些PMFs。我们在五种LLM上针对消费者民族中心主义进行案例研究,证明该方法能识别出聚合基准测试所掩盖的系统性国籍来源偏见。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。