arXiv:2508.00718cs.LG2025-08被引 2

开源工具包生成高质量表格数据,兼顾隐私与公平性。

Democratizing Tabular Data Access with an Open$\unicode{x2013}$Source Synthetic$\unicode{x2013}$Data SDK

  • 基于自回归框架的合成数据生成方法
  • 支持多表和时序数据,速度优于同类方案
  • 适合需要安全数据共享的研究与企业

机器学习发展严重依赖高质量数据,但隐私、商业利益和伦理问题导致数据获取日益困难。合成数据提供了一种可行方案,可在不泄露敏感信息的前提下实现数据的安全广泛使用。本文介绍MOSTLY AI合成数据SDK,一个专为生成高质量表格数据设计的开源工具包。该工具包集成差分隐私保障、公平性感知的数据生成及自动化质量评估功能,通过灵活易用的Python接口实现。基于TabularARGN自回归框架,支持多种数据类型及复杂多表和时序数据集,在性能上表现优异,尤其在速度和易用性方面有显著提升。目前该工具包已以云服务和本地安装两种方式部署,获得快速采用,展现出解决实际数据瓶颈的潜力,推动数据民主化进程。

原文摘要 · Abstract (English)

Machine learning development critically depends on access to high-quality data. However, increasing restrictions due to privacy, proprietary interests, and ethical concerns have created significant barriers to data accessibility. Synthetic data offers a viable solution by enabling safe, broad data usage without compromising sensitive information. This paper presents the MOSTLY AI Synthetic Data Software Development Kit (SDK), an open-source toolkit designed specifically for synthesizing high-quality tabular data. The SDK integrates robust features such as differential privacy guarantees, fairness-aware data generation, and automated quality assurance into a flexible and accessible Python interface. Leveraging the TabularARGN autoregressive framework, the SDK supports diverse data types and complex multi-table and sequential datasets, delivering competitive performance with notable improvements in speed and usability. Currently deployed both as a cloud service and locally installable software, the SDK has seen rapid adoption, highlighting its practicality in addressing real-world data bottlenecks and promoting widespread data democratization.

合成数据表格数据隐私保护开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。