arXiv:2501.15351cs.CYcs.LG2025-01被引 3

检测大模型在不同国家调查中的偏见,揭示其不公平表现

Fairness in LLM-Generated Surveys

  • 对比中美公众调查数据,评估大模型预测公平性
  • 美国数据上模型表现更好,源于训练数据的美国中心倾向
  • 不同国家的关键影响因素各异,如政治身份或教育水平

大型语言模型在文本生成与理解方面表现出色,尤其擅长模拟社会政治和经济模式,可作为传统调查的替代方案。然而,其在全球范围内的适用性仍存疑,因跨社会人口与地理背景的潜在偏见尚未被充分探索。本研究通过分析智利和美国的公开调查数据,考察大模型在不同人群中的表现,重点关注预测准确性和公平性指标。结果表明,模型在美式数据集上持续优于智利数据集,这种偏差源于以美国为中心的训练数据,即便控制了社会人口差异后依然显著。在美国,政治身份和种族显著影响预测准确性;而在智利,性别、教育程度和宗教信仰则起更重要作用。本文提出一种测量大模型社会人口偏见的新框架,为实现跨多元文化背景下的更公平、更公正模型性能提供路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel in text generation and understanding, especially in simulating socio-political and economic patterns, serving as an alternative to traditional surveys. However, their global applicability remains questionable due to unexplored biases across socio-demographic and geographic contexts. This study examines how LLMs perform across diverse populations by analyzing public surveys from Chile and the United States, focusing on predictive accuracy and fairness metrics. The results show performance disparities, with LLM consistently outperforming on U.S. datasets. This bias originates from the U.S.-centric training data, remaining evident after accounting for socio-demographic differences. In the U.S., political identity and race significantly influence prediction accuracy, while in Chile, gender, education, and religious affiliation play more pronounced roles. Our study presents a novel framework for measuring socio-demographic biases in LLMs, offering a path toward ensuring fairer and more equitable model performance across diverse socio-cultural contexts.

大模型偏见公平性评估社会调查

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。