发现大模型偏爱西方音乐,揭示了训练数据中的文化偏见。
Musical ethnocentrism in Large Language Models
- 通过列举顶级音乐人和评分各国音乐文化,检测模型偏好
- 模型在两国以上音乐贡献者中,西欧国家占比超70%
- 适合关注AI伦理与文化公平的研究者阅读
大型语言模型(LLMs)反映其训练数据中的偏见,而这些数据又源自创建者群体的价值判断。目前对地缘文化偏见的研究仍不足,这类偏见可能源于训练数据中不同地区与文化代表性不均,或隐含的价值判断。本文首次分析了大模型中的音乐偏见,重点考察ChatGPT和Mixtral。实验一要求模型列出各类别“前100位”音乐贡献者,并分析其国籍分布;实验二则让模型对各国音乐文化的不同方面进行数值评分。结果显示,两个实验中模型均表现出对西方音乐文化的显著偏好。
原文摘要 · Abstract (English)
Large Language Models (LLMs) reflect the biases in their training data and, by extension, those of the people who created this training data. Detecting, analyzing, and mitigating such biases is becoming a focus of research. One type of bias that has been understudied so far are geocultural biases. Those can be caused by an imbalance in the representation of different geographic regions and cultures in the training data, but also by value judgments contained therein. In this paper, we make a first step towards analyzing musical biases in LLMs, particularly ChatGPT and Mixtral. We conduct two experiments. In the first, we prompt LLMs to provide lists of the "Top 100" musical contributors of various categories and analyze their countries of origin. In the second experiment, we ask the LLMs to numerically rate various aspects of the musical cultures of different countries. Our results indicate a strong preference of the LLMs for Western music cultures in both experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。