发现语音和音乐处理中梅尔尺度存在文化偏见,提出更公平的替代方案。
Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music
- 用可学习的频域表示替代传统梅尔尺度,实现自适应频率分解。
- 对声调语言的识别错误率比非声调语言高12.5个百分点,非西方音乐性能下降15.7%。
- 推荐使用CQT或ERB尺度,能显著降低跨文化差异,计算开销仅增1%。
现代音频系统普遍采用源自20世纪40年代西方心理声学研究的梅尔尺度表征,可能蕴含文化偏见,导致系统性能在不同文化背景下出现系统性差异。本文全面评估了音频前端中的跨文化偏见,对比梅尔尺度与可学习替代方案(LEAF、SincNet)及心理声学变体(ERB、Bark、CQT),覆盖11种语言的语音识别、6个音乐数据集和10个欧洲城市的声音场景分类。控制实验隔离前端影响,保持模型结构与训练协议一致。结果表明,梅尔尺度在声调语言上的词错误率(WER)为31.2%,非声调语言为18.7%(差距12.5%),非西方音乐的F1分数下降15.7%。替代方法显著缓解偏差:LEAF通过自适应频率分配将语音差距降低34%,CQT使音乐性能差距减少52%,ERB尺度过滤仅增加1%计算开销即降低31%偏差。我们还发布了FairAudioBench,支持跨文化评估,并证明自适应频域分解是实现公平音频处理的可行路径。研究揭示基础信号处理选择如何传递偏见,为构建包容性音频系统提供关键指导。
原文摘要 · Abstract (English)
Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation of cross-cultural bias in audio front-ends, comparing mel-scale features with learnable alternatives (LEAF, SincNet) and psychoacoustic variants (ERB, Bark, CQT) across speech recognition (11 languages), music analysis (6 collections), and European acoustic scene classification (10 European cities). Our controlled experiments isolate front-end contributions while holding architecture and training protocols minimal and constant. Results demonstrate that mel-scale features yield 31.2% WER for tonal languages compared to 18.7% for non-tonal languages (12.5% gap), and show 15.7% F1 degradation between Western and non-Western music. Alternative representations significantly reduce these disparities: LEAF reduces the speech gap by 34% through adaptive frequency allocation, CQT achieves 52% reduction in music performance gaps, and ERB-scale filtering cuts disparities by 31% with only 1% computational overhead. We also release FairAudioBench, enabling cross-cultural evaluation, and demonstrate that adaptive frequency decomposition offers practical paths toward equitable audio processing. These findings reveal how foundational signal processing choices propagate bias, providing crucial guidance for developing inclusive audio systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。