检测大模型在情感识别中的性别偏见,发现仅靠提示工程无法有效缓解。
Gender Bias in Emotion Recognition by Large Language Models
- 通过人物描述判断情绪,测试大模型的性别偏见
- 基于训练的干预比提示工程更有效降低偏见
- 提醒开发者关注模型公平性,尤其在情感理解场景
大语言模型(LLMs)的快速发展及其在日常生活中的广泛应用,凸显了评估和保障其公平性的紧迫性。本文聚焦于情感心智理论领域,研究当给定一个人物描述及其环境时,大模型在回答‘这个人感觉如何?’这一问题时是否表现出性别偏见。我们提出了多种去偏策略并进行评估,结果表明,仅依赖推理阶段的提示工程等方法难以显著减少偏见;真正有效的减偏需要基于训练阶段的干预措施。本研究强调了在情感理解任务中对模型公平性进行系统性评估与改进的重要性。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) and their growing integration into daily life underscore the importance of evaluating and ensuring their fairness. In this work, we examine fairness within the domain of emotional theory of mind, investigating whether LLMs exhibit gender biases when presented with a description of a person and their environment and asked, ''How does this person feel?''. Furthermore, we propose and evaluate several debiasing strategies, demonstrating that achieving meaningful reductions in bias requires training based interventions rather than relying solely on inference-time prompt-based approaches such as prompt engineering, etc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。