聚焦NLP从业者数据公平实践,推动责任共担与治理改革
Advancing Data Equity: Practitioner Responsibility and Accountability in NLP Data Practices
- 从从业者视角切入,分析数据公平认知与实践困境
- 发现商业目标与公平承诺间存在持续张力
- 倡导参与式数据流程与结构化治理改革
尽管研究多关注算法偏见的识别与审计以保障公平的AI发展,但对直接参与数据集构建、标注与部署的NLP从业者如何感知和应对数据公平问题仍了解不足。本研究首次聚焦从业者视角,将其经验与多尺度人工智能治理框架相联系,提出跨技术、政策与社区领域的参与式建议。基于2024年问卷调查与焦点小组访谈,我们探讨了美国本土NLP数据从业者如何理解公平性,如何应对组织与系统性约束,并参与如《美国人工智能权利法案》等新兴治理倡议。研究发现,商业目标与公平承诺之间存在持续张力,且普遍呼吁建立更具参与性和问责性的数据工作流程。我们批判性回应数据多样性与‘多样性洗白’的争论,主张提升NLP公平性需依赖支持从业者自主权与社区同意的结构性治理改革。
原文摘要 · Abstract (English)
While research has focused on surfacing and auditing algorithmic bias to ensure equitable AI development, less is known about how NLP practitioners - those directly involved in dataset development, annotation, and deployment - perceive and navigate issues of NLP data equity. This study is among the first to center practitioners' perspectives, linking their experiences to a multi-scalar AI governance framework and advancing participatory recommendations that bridge technical, policy, and community domains. Drawing on a 2024 questionnaire and focus group, we examine how U.S.-based NLP data practitioners conceptualize fairness, contend with organizational and systemic constraints, and engage emerging governance efforts such as the U.S. AI Bill of Rights. Findings reveal persistent tensions between commercial objectives and equity commitments, alongside calls for more participatory and accountable data workflows. We critically engage debates on data diversity and diversity washing, arguing that improving NLP equity requires structural governance reforms that support practitioner agency and community consent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。