arXiv:2603.22260cs.CL2026-03被引 2

语音接口让AI更易用,却可能加剧性别偏见。

Greater accessibility can amplify discrimination in generative AI

  • 语音输入使模型根据声音自动判断性别并产生刻板印象
  • 语音交互的偏见程度高于文本,且1000人调查证实用户反感
  • 调节语音音高可减轻歧视,适合政策制定者和伦理设计者

数亿人依赖大语言模型进行教育、工作甚至医疗。然而这些模型会复制并放大训练数据中的社会偏见。文本界面对读写能力弱、有运动障碍或仅使用手机的用户构成障碍。语音交互虽能提升可及性,但语音自带身份线索,用户难以隐藏,引发公平性担忧。我们发现,支持语音的LLM会基于说话者声音系统性地偏向性别刻板印象,如职业和形容词选择,其偏见程度超过文本交互。补充调查显示,1000名用户中不常使用聊天机器人的群体最反感属性推断,一旦知晓会更易放弃使用。实验表明,调节语音音高可有效控制性别歧视输出。研究揭示:通过语音扩展可及性的同时,可能引入新的歧视路径,需将公平与可及性同步考虑。

原文摘要 · Abstract (English)

Hundreds of millions of people rely on large language models (LLMs) for education, work, and even healthcare. Yet these models are known to reproduce and amplify social biases present in their training data. Moreover, text-based interfaces remain a barrier for many, for example, users with limited literacy, motor impairments, or mobile-only devices. Voice interaction promises to expand accessibility, but unlike text, speech carries identity cues that users cannot easily mask, raising concerns about whether accessibility gains may come at the cost of equitable treatment. Here we show that audio-enabled LLMs exhibit systematic gender discrimination, shifting responses toward gender-stereotyped adjectives and occupations solely on the basis of speaker voice, and amplifying bias beyond that observed in text-based interaction. Thus, voice interfaces do not merely extend text models to a new modality but introduce distinct bias mechanisms tied to paralinguistic cues. Complementary survey evidence ($n=1,000$) shows that infrequent chatbot users are most hesitant to undisclosed attribute inference and most likely to disengage when such practices are revealed. To demonstrate a potential mitigation strategy, we show that pitch manipulation can systematically regulate gender-discriminatory outputs. Overall, our findings reveal a critical tension in AI development: efforts to expand accessibility through voice interfaces simultaneously create new pathways for discrimination, demanding that fairness and accessibility be addressed in tandem.

语音交互性别偏见公平性可及性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。