arXiv:2503.03506cs.LGcs.AI2025-03被引 1

从隐私角度重构合成数据分类,助力监管与应用。

Opinion: Revisiting synthetic data classifications from a privacy perspective

  • 按生成方法与数据源联合划分合成数据类型。
  • 新分类框架支持深度生成技术等新技术发展。
  • 适合政策制定者与数据合规团队参考使用。

合成数据正成为满足人工智能发展日益增长数据需求的低成本解决方案,其生成方式或基于现有知识,或源自真实数据。传统将合成数据分为混合、部分或完全合成的分类方法已难以反映当前多样化的生成技术。生成方法及其数据来源共同决定了合成数据的特性,进而影响其实际应用场景。本文主张一种新的合成数据分类方式,更贴合隐私视角,以促进合成数据生成与处理的监管指引。该分类框架具有灵活性,能适应深度生成等新兴技术,并为未来应用提供更实用的指导。

原文摘要 · Abstract (English)

Synthetic data is emerging as a cost-effective solution necessary to meet the increasing data demands of AI development, created either from existing knowledge or derived from real data. The traditional classification of synthetic data types into hybrid, partial or fully synthetic datasets has limited value and does not reflect the ever-increasing methods to generate synthetic data. The generation method and their source jointly shape the characteristics of synthetic data, which in turn determines its practical applications. We make a case for an alternative approach to grouping synthetic data types that better reflect privacy perspectives in order to facilitate regulatory guidance in the generation and processing of synthetic data. This approach to classification provides flexibility to new advancements like deep generative methods and offers a more practical framework for future applications.

合成数据隐私保护分类框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。