arXiv:2604.15990cs.CYcs.AI2026-04

警惕数据保护反成新剥削,揭示算法如何加剧弱势群体风险

From Vulnerable Data Subjects to Vulnerabilizing Data Practices: Navigating the Protection Paradox in AI-Based Analyses of Platformized Lives

  • 从静态脆弱性转向分析数据实践如何主动制造脆弱
  • 实证发现保护性研究可能引发计算暴露与过度简化
  • 提出四阶段伦理框架,指导研究人员规避数据剥削

本文探讨了从将脆弱性视为数据主体固有属性,转向审视其如何通过数据实践被主动建构的范式转变。不同于聚焦缺失或反向数据的反思性伦理框架,本文关注平台化生活中的数据丰沛状态——即海量数据已存在,研究者的挑战在于如何处理这一既存数据流。我们主张,数据科学的伦理完整性不仅关乎研究对象,更取决于技术流程如何将“脆弱个体”转化为可进一步被边缘化的数据主体。通过一个AI for Social Good案例:记者请求利用计算机视觉量化商业化的YouTube‘家庭视频博客’中儿童出现情况以推动监管,揭示了‘保护悖论’——旨在保护弱势群体的数据行动可能意外导致新的计算暴露、还原主义和数据攫取。通过对该请求的技术管道进行方法学解构,本文展示微小技术决策如何具有伦理决定性。进而提出一套反思性伦理协议,围绕数据集设计、操作化、推断和传播四个关键节点,识别出技术问题与伦理张力,提供具体提示以应对四种交叉性的致脆弱因素:暴露、货币化、叙事固化和算法优化。

原文摘要 · Abstract (English)

This paper traces a conceptual shift from understanding vulnerability as a static, essentialized property of data subjects to examining how it is actively enacted through data practices. Unlike reflexive ethical frameworks focused on missing or counter-data, we address the condition of abundance inherent to platformized life-a context where a near inexhaustible mass of data points already exists, shifting the ethical challenge to the researcher's choices in operating upon this existing mass. We argue that the ethical integrity of data science depends not just on who is studied, but on how technical pipelines transform "vulnerable" individuals into data subjects whose vulnerability can be further precarized. We develop this argument through an AI for Social Good (AI4SG) case: a journalist's request to use computer vision to quantify child presence in monetized YouTube 'family vlogs' for regulatory advocacy. This case reveals a "protection paradox": how data-driven efforts to protect vulnerable subjects can inadvertently impose new forms of computational exposure, reductionism, and extraction. Using this request as a point of departure, we perform a methodological deconstruction of the AI pipeline to show how granular technical decisions are ethically constitutive. We contribute a reflexive ethics protocol that translates these insights into a reflexive roadmap for research ethics surrounding platformized data subjects. Organized around four critical junctures-dataset design, operationalization, inference, and dissemination-the protocol identifies technical questions and ethical tensions where well-intentioned work can slide into renewed extraction or exposure. For every decision point, the protocol offers specific prompts to navigate four cross-cutting vulnerabilizing factors: exposure, monetization, narrative fixing, and algorithmic optimization. Rather than uncritically...

算法伦理数据剥削平台治理社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。