ASR系统对少数方言用户的情感伤害远超错误率,需重思用户体验。
"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias
- 通过四地用户研究,揭示方言用户为适配系统付出隐形语言劳动。
- 多数用户因系统不兼容而产生挫败感与自我怀疑,情绪代价被忽视。
- 强调评估应超越准确率,纳入情感负担与文化认同维度,适合人机交互研究者。
针对自动语音识别(ASR)中的偏见问题,现有研究多聚焦未代表方言群体的错误率,却较少关注系统偏差对用户的实际影响:系统失败如何塑造用户的生活体验、引发何种情绪反应,以及持续失效带来的心理代价?我们在美国四个不同英语方言区(亚特兰大、墨西哥湾沿岸、迈阿密海滩、图森)开展用户研究。结果显示,大多数参与者认为技术未考虑其文化背景,需不断调整才能实现基本功能。尽管如此,他们仍对ASR性能抱有高期待,并愿意参与模型改进。开放式叙事分析显示,用户不仅感到沮丧、恼怒和能力不足,更在长期使用中内化失败为个人缺陷。他们主动进行语码转换、过度清晰发音及情绪管理以适应系统,但其语言文化知识未被技术承认。当前以准确率为单一标准的公平性评估,忽略了算法排斥所造成的隐性情感劳动、持续自我监控的认知负荷,以及在母语中感到不适的心理代价。
原文摘要 · Abstract (English)
Studies on bias in Automatic Speech Recognition (ASR) tend to focus on reporting error rates for speakers of underrepresented dialects, yet less research examines the human side of system bias: how do system failures shape users' lived experiences, how do users feel about and react to them, and what emotional toll do these repeated failures exact? We conducted user experience studies across four U.S. locations (Atlanta, Gulf Coast, Miami Beach, and Tucson) representing distinct English dialect communities. Our findings reveal that most participants report technologies fail to consider their cultural backgrounds and require constant adjustment to achieve basic functionality. Despite these experiences, participants maintain high expectations for ASR performance and express strong willingness to contribute to model improvement. Qualitative analysis of open-ended narratives exposes the deeper costs of these failures. Participants report frustration, annoyance, and feelings of inadequacy, yet the emotional impact extends beyond momentary reactions. Participants recognize that systems were not designed for them, yet often internalize failures as personal inadequacy despite this critical awareness. They perform extensive invisible labor, including code-switching, hyper-articulation, and emotional management, to make failing systems functional. Meanwhile, their linguistic and cultural knowledge remains unrecognized by technologies that encode particular varieties as standard while rendering others marginal. These findings demonstrate that algorithmic fairness assessments based on accuracy metrics alone miss critical dimensions of harm: the emotional labor of managing repeated technological rejection, the cognitive burden of constant self-monitoring, and the psychological toll of feeling inadequate in one's native language variety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。