arXiv:2602.24002cs.CL2026-02被引 1

检测YouTube西班牙语字幕系统对不同方言和性别的识别偏差

Dialect and Gender Bias in YouTube's Spanish Captioning System

  • 对比不同地区西班牙语方言的自动字幕准确率
  • 发现女性发言者及特定方言在字幕生成中表现更差
  • 提醒平台需考虑语言多样性以减少算法偏见

西班牙语是21个国家的官方语言,使用者超过4.41亿人。由于地域差异,西班牙语存在多种方言。视频平台如YouTube依赖自动语音识别系统生成西班牙语字幕,但仅提供单一字幕选项。本研究分析该系统在不同西班牙语方言中的表现,通过比较来自不同地区的男女发言者字幕质量,发现系统存在系统性偏差,可归因于特定方言。结果表明,数字平台部署的算法技术必须针对用户群体的语言多样性进行校准。

原文摘要 · Abstract (English)

Spanish is the official language of twenty-one countries and is spoken by over 441 million people. Naturally, there are many variations in how Spanish is spoken across these countries. Media platforms such as YouTube rely on automatic speech recognition systems to make their content accessible to different groups of users. However, YouTube offers only one option for automatically generating captions in Spanish. This raises the question: could this captioning system be biased against certain Spanish dialects? This study examines the potential biases in YouTube's automatic captioning system by analyzing its performance across various Spanish dialects. By comparing the quality of captions for female and male speakers from different regions, we identify systematic disparities which can be attributed to specific dialects. Our study provides further evidence that algorithmic technologies deployed on digital platforms need to be calibrated to the diverse needs and experiences of their user populations.

语音识别语言偏见算法公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。