arXiv:2604.02660cs.CL2026-04被引 1

构建模板框架评估大模型中的阶级偏见,发现生活判断偏见是教育决策的10倍。

SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models

  • 基于模板生成240个任务,覆盖8大主题18类议题,系统评估阶级偏见。
  • 13个主流大模型在3120次回答中偏见率差异显著,最高达33.75%。
  • 揭示生活类判断偏见远高于教育类,现有防护对特定刻板印象无效。

随着大语言模型(LLMs)在关键领域决策系统中的广泛应用,理解并缓解其偏见已成为负责任AI部署的必要前提。尽管针对种族、性别等属性的偏见评估框架已大量出现,但社会经济地位偏见因现实影响广泛却仍严重缺乏研究。本文提出SocioEval,一个基于模板的框架,通过决策任务系统评估基础模型中的社会经济偏见。该框架包含8个主题和18个议题,生成240个提示,涵盖6种类别组合。我们在3,120次响应上评估了13个前沿大模型,采用严格的三阶段标注协议,发现偏见率差异显著(0.42%–33.75%)。结果表明,偏见在不同主题中表现各异:生活方式判断的偏见是教育相关决策的10倍;现有部署防护能有效防止显性歧视,但对特定领域刻板印象表现出脆弱性。SocioEval为审计语言模型中的阶级偏见提供了可扩展、可拓展的基础。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) increasingly power decision-making systems across critical domains, understanding and mitigating their biases becomes essential for responsible AI deployment. Although bias assessment frameworks have proliferated for attributes such as race and gender, socioeconomic status bias remains significantly underexplored despite its widespread implications in the real world. We introduce SocioEval, a template-based framework for systematically evaluating socioeconomic bias in foundation models through decision-making tasks. Our hierarchical framework encompasses 8 themes and 18 topics, generating 240 prompts across 6 class-pair combinations. We evaluated 13 frontier LLMs on 3,120 responses using a rigorous three-stage annotation protocol, revealing substantial variation in bias rates (0.42\%-33.75\%). Our findings demonstrate that bias manifests differently across themes lifestyle judgments show 10$\times$ higher bias than education-related decisions and that deployment safeguards effectively prevent explicit discrimination but show brittleness to domain-specific stereotypes. SocioEval provides a scalable, extensible foundation for auditing class-based bias in language models.

大模型偏见社会经济评估框架伦理AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。