审计主流音频数据集,发现性别偏见与版权问题。
Sound Check: Auditing Audio Datasets
- 筛选7个主流音频数据集进行深度审计。
- 发现数据集存在性别偏见、歧视性刻板印象和大量版权内容。
- 开发网页工具供艺术家查询自身作品是否被收录。
生成式音频模型正快速进步并广泛使用——多个强大模型已开放权重,部分科技公司也推出了高质量生成音频产品。然而,尽管已有研究指出生成式视觉和文本模型训练数据中存在的诸多伦理问题,我们对生成式音频数据集的类似问题了解甚少,包括偏见、毒性及知识产权问题。为填补这一空白,我们对数百个音频数据集进行了文献综述,并选取其中7个最具代表性者进行深入审计。结果发现,这些数据集存在针对女性的偏见,包含对边缘化群体的有害刻板印象,且包含大量受版权保护的内容。为帮助艺术家判断其作品是否被纳入主流音频数据集,并促进对数据集内容的探索,我们开发了网页工具 audio datasets exploration tool,网址为 https://audio-audit.vercel.app。
原文摘要 · Abstract (English)
Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio products. Yet, while prior work has enumerated many ethical issues stemming from the data on which generative visual and textual models have been trained, we have little understanding of similar issues with generative audio datasets, including those related to bias, toxicity, and intellectual property. To bridge this gap, we conducted a literature review of hundreds of audio datasets and selected seven of the most prominent to audit in more detail. We found that these datasets are biased against women, contain toxic stereotypes about marginalized communities, and contain significant amounts of copyrighted work. To enable artists to see if they are in popular audio datasets and facilitate exploration of the contents of these datasets, we developed a web tool audio datasets exploration tool at https://audio-audit.vercel.app.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。