分析手语模型偏见并提出有效缓解方法,提升技术公平性。
Studying and Mitigating Biases in Sign Language Understanding Models
- 利用ASL Citizen数据集的用户信息研究手语模型偏见来源。
- 多种缓解策略降低性能差异,且不损害整体准确率。
- 公开参与者人口统计信息,推动后续公平性研究。
确保手语技术的益处公平惠及所有社区成员至关重要。因此,需关注设计或使用这些资源时可能引发的偏见与不平等。众包手语数据集(如ASL Citizen)是提升可及性和保护语言多样性的宝贵资源,但必须谨慎使用以避免强化既有偏见。本文利用ASL Citizen数据集中丰富的参与者人口统计和词汇特征信息,系统研究并记录了基于众包数据训练的手语模型可能产生的偏见。进一步地,在模型训练中应用多种偏见缓解技术,发现这些方法能有效减少性能差异,同时保持准确率不变。本文发布后,将公开ASL Citizen数据集中参与者的匿名人口统计信息,以促进该领域未来的偏见缓解研究。
原文摘要 · Abstract (English)
Ensuring that the benefits of sign language technologies are distributed equitably among all community members is crucial. Thus, it is important to address potential biases and inequities that may arise from the design or use of these resources. Crowd-sourced sign language datasets, such as the ASL Citizen dataset, are great resources for improving accessibility and preserving linguistic diversity, but they must be used thoughtfully to avoid reinforcing existing biases. In this work, we utilize the rich information about participant demographics and lexical features present in the ASL Citizen dataset to study and document the biases that may result from models trained on crowd-sourced sign datasets. Further, we apply several bias mitigation techniques during model training, and find that these techniques reduce performance disparities without decreasing accuracy. With the publication of this work, we release the demographic information about the participants in the ASL Citizen dataset to encourage future bias mitigation work in this space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。