检测合成语音服务中的口音偏见,揭示技术不平等对数字排斥的影响
"It's not a representation of me": Examining Accent Bias and Digital Exclusion in Synthetic AI Voice Services
- 通过问卷与访谈结合,评估五种英语口音在两个语音服务中的表现差异
- 发现不同口音的语音生成质量存在明显差距,口音弱势者体验更差
- 呼吁开发者与政策制定者关注语音技术的公平性,防止加剧语言歧视
近年来,人工智能语音生成与语音克隆技术已实现自然流畅的语音合成与精准的声音复制,但其在多元口音与语言特征下的社会技术影响尚未充分理解。本研究采用混合方法,通过问卷与访谈评估Speechify和ElevenLabs两款合成语音服务的技术表现,并探索用户真实经历如何影响其对口音差异的认知。结果揭示五种英语地区口音在技术性能上存在显著差异,当前语音生成技术可能无意中强化语言特权与口音歧视,进而导致新的数字排斥形式。研究强调需推动包容性设计与监管,为开发者、政策制定者及组织提供可操作的洞察,确保语音AI技术的公平与社会责任。
原文摘要 · Abstract (English)
Recent advances in artificial intelligence (AI) speech generation and voice cloning technologies have produced naturalistic speech and accurate voice replication, yet their influence on sociotechnical systems across diverse accents and linguistic traits is not fully understood. This study evaluates two synthetic AI voice services (Speechify and ElevenLabs) through a mixed methods approach using surveys and interviews to assess technical performance and uncover how users' lived experiences influence their perceptions of accent variations in these speech technologies. Our findings reveal technical performance disparities across five regional, English-language accents and demonstrate how current speech generation technologies may inadvertently reinforce linguistic privilege and accent-based discrimination, potentially creating new forms of digital exclusion. Overall, our study highlights the need for inclusive design and regulation by providing actionable insights for developers, policymakers, and organizations to ensure equitable and socially responsible AI speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。