arXiv:2601.02858cs.CL2026-01

用反向提示法评估大模型文化对齐能力,发现生成比判断更难。

To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs

  • 用反向提示让模型从行为中判断用户身份,避免生成任务的偏差干扰
  • 真实行为判断准确率高于模拟行为,但个体层面两者差距缩小
  • 揭示个性化推荐的极限,适合研究模型偏见与评估方法论

社会人口学提示(SDP)通过人口统计学代理提示大语言模型生成文化相关输出,常暴露模型刻板印象和偏见。然而,该方法易受提示敏感性、解码参数及生成任务本身复杂性等混淆因素影响,难以区分性能不佳是源于偏见还是任务设计缺陷。为解决此问题,本文采用逆向社会人口学提示(ISDP),即让模型根据真实与模拟用户行为判断其人口统计属性。基于Goodreads-CSI数据集(Saha et al., 2025),该数据捕捉印度、墨西哥和美国用户理解英文书评的难度,测试了Aya-23、Gemma-2、GPT-4o和LLaMA-3.1四款模型。结果表明,模型在真实行为上的表现优于模拟行为,但在个体层面,两类表现均下降并趋于接近,说明个性化存在固有局限。

原文摘要 · Abstract (English)

Socio-demographic prompting (SDP) - prompting Large Language Models (LLMs) using demographic proxies to generate culturally aligned outputs - often shows LLM responses as stereotypical and biased. While effective in assessing LLMs' cultural competency, SDP is prone to confounding factors such as prompt sensitivity, decoding parameters, and the inherent difficulty of generation over discrimination tasks due to larger output spaces. These factors complicate interpretation, making it difficult to determine if the poor performance is due to bias or the task design. To address this, we use inverse socio-demographic prompting (ISDP), where we prompt LLMs to discriminate and predict the demographic proxy from actual and simulated user behavior from different users. We use the Goodreads-CSI dataset (Saha et al., 2025), which captures difficulty in understanding English book reviews for users from India, Mexico, and the USA, and test four LLMs: Aya-23, Gemma-2, GPT-4o, and LLaMA-3.1 with ISDP. Results show that models perform better with actual behaviors than simulated ones, contrary to what SDP suggests. However, performance with both behavior types diminishes and becomes nearly equal at the individual level, indicating limits to personalization.

大模型评估文化对齐提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。