arXiv:2503.13149cs.AIcs.CL2025-03被引 1

用心理测量学方法检测大模型隐性偏见,发现多数回避立场而非有倾向。

Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs

  • 基于项目反应理论建模,量化回答回避与潜在偏见
  • 实测主流模型多回避政治立场而非展现偏见
  • 为公平AI治理提供可量化的对齐工具

我们提出一种基于项目反应理论(IRT)的框架,无需依赖主观人类判断即可检测和量化大语言模型(LLMs)中的社会经济偏见。与传统方法不同,IRT 能够考虑题目难度,从而更准确地估计意识形态偏见。我们微调了两个 LLM 家族(Meta-LLaMa 3.2-1B-Instruct 与 ChatGPT 3.5)以代表不同意识形态立场,并采用两阶段方法:(1) 建模回答回避行为;(2) 估算已回答内容中的感知偏见。结果表明,现成的 LLM 通常避免意识形态互动,而非表现出偏见,挑战了先前关于其党派性的观点。该经实证验证的框架增强了 AI 对齐研究,并推动更公平的 AI 治理。

原文摘要 · Abstract (English)

We introduce an Item Response Theory (IRT)-based framework to detect and quantify socioeconomic bias in large language models (LLMs) without relying on subjective human judgments. Unlike traditional methods, IRT accounts for item difficulty, improving ideological bias estimation. We fine-tune two LLM families (Meta-LLaMa 3.2-1B-Instruct and Chat- GPT 3.5) to represent distinct ideological positions and introduce a two-stage approach: (1) modeling response avoidance and (2) estimating perceived bias in answered responses. Our results show that off-the-shelf LLMs often avoid ideological engagement rather than exhibit bias, challenging prior claims of partisanship. This empirically validated framework enhances AI alignment research and promotes fairer AI governance.

大模型偏见心理测量AI对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。