自动情绪识别对男性表达的情绪检测普遍不准,存在系统性偏差。
Automatic Classifiers Underdetect Emotions Expressed by Men
- 用百万级自标注文本数据,检验414种模型与情绪类别的组合
- 男性文本的情绪识别错误率显著高于女性,各模型均存在此现象
- 提醒使用时需警惕性别未知或变化样本,尤其大模型更需谨慎
自动情感与情绪分类器的广泛应用使得其在不同人群中的可靠性至关重要。然而,现有评估多依赖第三方标注,可能掩盖系统性偏差。本文利用超过一百万条自标注帖子的大型数据集和预注册研究设计,考察414种模型与情绪类别组合中的性别偏差。结果表明,无论模型类型或情绪类别,男性作者文本的情绪识别错误率始终高于女性。我们量化了该偏差对下游应用的影响,指出当前机器学习工具(包括大语言模型)在样本性别组成未知或可变时应谨慎使用。研究证明,情感分析仍未解决,尤其在跨人口群体的公平性方面仍存挑战。
原文摘要 · Abstract (English)
The widespread adoption of automatic sentiment and emotion classifiers makes it important to ensure that these tools perform reliably across different populations. Yet their reliability is typically assessed using benchmarks that rely on third-party annotators rather than the individuals experiencing the emotions themselves, potentially concealing systematic biases. In this paper, we use a unique, large-scale dataset of more than one million self-annotated posts and a pre-registered research design to investigate gender biases in emotion detection across 414 combinations of models and emotion-related classes. We find that across different types of automatic classifiers and various underlying emotions, error rates are consistently higher for texts authored by men compared to those authored by women. We quantify how this bias could affect results in downstream applications and show that current machine learning tools, including large language models, should be applied with caution when the gender composition of a sample is not known or variable. Our findings demonstrate that sentiment analysis is not yet a solved problem, especially in ensuring equitable model behaviour across demographic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。