arXiv:2508.09232cs.MMcs.AI2025-08AAAI被引 3

为社交媒体数据研究设计隐私优先的合规流程,解决法律与实践脱节问题。

PETLP: A Privacy-by-Design Pipeline for Social Media Data in AI Research

  • 将数据保护评估融入数据处理全流程,动态更新合规策略。
  • 揭示科研机构可依据欧盟DSM第3条突破平台限制,商业主体则受限于服务条款。
  • 指出社交媒体数据无法真正匿名,强调模型分发存在法律模糊地带。

社交媒体数据使人工智能研究面临欧盟GDPR、版权法及平台条款的多重合规义务,但现有框架未能整合这些监管领域,导致研究人员缺乏统一指引。本文提出隐私优先的PETLP(Privacy-by-design Extract, Transform, Load, and Present)框架,将法律保障嵌入扩展的ETL流程中。核心在于将数据保护影响评估视为动态演进的文档,从研究注册到成果发布持续更新。通过系统分析Reddit数据,我们发现科研机构可依据DSM第3条规避平台限制,而商业实体必须遵守服务条款,但GDPR义务对所有主体均适用。研究证实社交媒体数据无法实现真正匿名,并揭示了数据集创建与模型分发之间的法律空白。通过将合规决策转化为可操作工作流并简化机构数据管理方案,PETLP帮助研究人员在复杂法规中自信前行,弥合法律要求与研究实践间的鸿沟。

原文摘要 · Abstract (English)

Social media data presents AI researchers with overlapping obligations under the GDPR, copyright law, and platform terms -- yet existing frameworks fail to integrate these regulatory domains, leaving researchers without unified guidance. We introduce PETLP (Privacy-by-design Extract, Transform, Load, and Present), a compliance framework that embeds legal safeguards directly into extended ETL pipelines. Central to PETLP is treating Data Protection Impact Assessments as living documents that evolve from pre-registration through dissemination. Through systematic Reddit analysis, we demonstrate how extraction rights fundamentally differ between qualifying research organisations (who can invoke DSM Article 3 to override platform restrictions) and commercial entities (bound by terms of service), whilst GDPR obligations apply universally. We demonstrate why true anonymisation remains unachievable for social media data and expose the legal gap between permitted dataset creation and uncertain model distribution. By structuring compliance decisions into practical workflows and simplifying institutional data management plans, PETLP enables researchers to navigate regulatory complexity with confidence, bridging the gap between legal requirements and research practice.

隐私保护合规框架数据伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。