用间接测试评估大模型在能力、社交、道德上的隐性偏见
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
- 将刻板印象重构为心智理论的多维失败,设计词联想与情感归因任务
- 在8个顶尖大模型上发现复杂偏见结构,验证框架有效性
- 适合关注AI伦理、偏见检测的研究者与开发者
大型语言模型(LLMs)的心智理论(ToM)指其推断他人心理状态的能力,该能力的缺失常表现为系统性隐性偏见。传统直接提问方法易遭拒绝且难以捕捉其细微多维特性。为此,我们提出MIST,将刻板印象内容重构为能力、社交性与道德性三个维度的ToM失败。该框架引入两项间接任务:词联想偏见测试(WABT)评估隐性词汇关联,情感归因测试(AAT)测量隐性情绪倾向,旨在不触发模型回避的情况下揭示潜在刻板印象。在八个先进大模型上的广泛实验表明,该框架能有效揭示复杂的偏见结构并具备更高鲁棒性。所有数据与代码将公开。
原文摘要 · Abstract (English)
Theory of Mind (ToM) in Large Language Models (LLMs) refers to the model's ability to infer the mental states of others, with failures in this ability often manifesting as systemic implicit biases. Assessing this challenge is difficult, as traditional direct inquiry methods are often met with refusal to answer and fail to capture its subtle and multidimensional nature. Therefore, we propose MIST, which reconceptualizes the content model of stereotypes into multidimensional failures of ToM, specifically in the domains of competence, sociability, and morality. The framework introduces two indirect tasks. The Word Association Bias Test (WABT) assesses implicit lexical associations, while the Affective Attribution Test (AAT) measures implicit emotional tendencies, aiming to uncover latent stereotypes without triggering model avoidance. Through extensive experimentation on eight state-of-the-art LLMs, our framework demonstrates the ability to reveal complex bias structures and improved robustness. All data and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。