通过幽默测试大模型对不同身份的不公平反应,发现特权者说笑更易被拒绝。
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor

- 用身份互换实验观察模型对幽默的响应差异。
- 特权者说笑被拒率高67.5%,恶意判定多64.7%。
- 揭示模型中敏感性与刻板印象并存的公平性困境。
幽默反映社会认知:我们觉得好笑的内容往往映射出自身及对他人的评判标准。当语言模型处理幽默时,其反应暴露了训练数据中内化的社会假设。本文通过固定其他因素、仅交换说话人与对象身份,研究基于幽默的反事实不公平性。框架涵盖三类任务:幽默生成拒绝、说话人意图推断、关系/社会影响预测,覆盖无身份指向与特定身份贬损类幽默。引入可解释的偏见度量,捕捉身份互换下的不对称模式。在多个前沿模型上的实验显示,持续存在关系性差异:由特权者说出的笑话被拒绝比例高出67.5%,被判定为恶意的概率高64.7%,在5分制的社会伤害评分中平均高1.5分。这些模式表明,生成式模型中敏感性与刻板印象共存,加剧了公平性与文化契合的挑战。
原文摘要 · Abstract (English)
Humor holds up a mirror to social perception: what we find funny often reflects who we are and how we judge others. When language models engage with humor, their reactions expose the social assumptions they have internalized from training data. In this paper, we investigate counterfactual unfairness through humor by observing how the model's responses change when we swap who speaks and who is addressed while holding other factors constant. Our framework spans three tasks: humor generation refusal, speaker intention inference, and relational/societal impact prediction, covering both identity-agnostic humor and identity-specific disparagement humor. We introduce interpretable bias metrics that capture asymmetric patterns under identity swaps. Experiments across state-of-the-art models reveal consistent relational disparities: jokes told by privileged speakers are refused up to 67.5% more often, judged as malicious 64.7% more frequently, and rated up to 1.5 points higher in social harm on a 5-point scale. These patterns highlight how sensitivity and stereotyping coexist in generative models, complicating efforts toward fairness and cultural alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。