研究大模型对乐器的性别偏见,发现文本中偏见最严重
Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

- 构建跨模态数据集,分析模型对22种乐器的性别关联
- 92%结果符合社会科学研究的性别刻板印象
- 音频模态偏见最弱,文本最强,提示模态差异影响偏见
大型语言模型(LLMs)广泛应用于日常信息获取,引发对其可能延续社会偏见的担忧。本研究从乐器的性别刻板印象出发,结合社会科学成果,提出 Symphony-Bias 多模态数据集,涵盖文本、视觉和音频三种模态。评估了十种具有不同架构与规模的多模态模型,针对22种乐器,分析其与{男性, 女性, 非二元}三类性别的关联。结果显示,92%的乐器层面结果与既有社会科学研究一致,其中竖琴与鼓在所有模型和模态中均表现出高度一致的性别关联。进一步发现,偏见程度在音频模态最弱,视觉较强,文本最强,表明模态特异性表征会差异化放大乐器与性别的关联。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments. Building on social-science research on the cultural gender-typing of instruments, we introduce Symphony-Bias, a parallel multimodal dataset spanning text, vision, and audio. We evaluate ten multimodal models with diverse architectures and scales across 22 musical instruments, analyzing how they associate each instrument with three gender categories: {male, female, non-binary}, across three modalities: {text, vision, audio}. Our results show that 92\% of instrument-level outcomes align with prior social-science findings, with the harp and drums showing particularly consistent gendered associations across all evaluated models and modalities. We further find that alignment with social stereotypes is weakest in audio, stronger in vision, and strongest in text, suggesting that modality-specific representations can differentially amplify gendered associations with musical instruments.\footnote{The Symphony-Bias dataset will be publicly released upon acceptance of the paper.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。