arXiv:2503.00093cs.CYcs.AI2025-03ICML被引 9

用社会科学方法改进大模型偏见探测,让结果更可信、可比较。

Rethinking LLM Bias Probing Using Lessons from the Social Sciences

  • 借鉴社会科学研究,建立系统化探测框架
  • 解决探针选择混乱和结果冲突问题
  • 帮助判断偏见是否影响真实用户行为

大模型偏见探测的普及带来了三大挑战:缺乏合理选择探针的标准、无法调和不同探针的矛盾结果、缺乏形式化框架来判断探测结果能否推广到真实用户行为。本文通过整合社会科学研究的实用洞见,系统化地重构大模型社会偏见探测。提出EcoLevels框架,用于确定合适的偏见探针、调和不同探针间的矛盾发现,并预测偏见的泛化能力。研究指出,许多大模型探针直接借鉴人类偏见研究,而社会科学领域早已面临类似挑战。因此,下一代大模型偏见探测应充分借鉴数十年的社会科学成果。

原文摘要 · Abstract (English)

The proliferation of LLM bias probes introduces three significant challenges: (1) we lack principled criteria for choosing appropriate probes, (2) we lack a system for reconciling conflicting results across probes, and (3) we lack formal frameworks for reasoning about when (and why) probe results will generalize to real user behavior. We address these challenges by systematizing LLM social bias probing using actionable insights from social sciences. We then introduce EcoLevels - a framework that helps (a) determine appropriate bias probes, (b) reconcile conflicting findings across probes, and (c) generate predictions about bias generalization. Overall, we ground our analysis in social science research because many LLM probes are direct applications of human probes, and these fields have faced similar challenges when studying social bias in humans. Based on our work, we suggest how the next generation of LLM bias probing can (and should) benefit from decades of social science research.

大模型偏见社会科学研究探针设计可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。