构建多维评估框架,量化合成数据的分布还原与隐私保护能力
Benchmarking Synthetic Tabular Data: A Multi-Dimensional Evaluation Framework
- 基于留出法设计评估流程,融合低维高维分布比对
- 通过嵌入相似度与近邻距离实现可解释质量诊断
- 支持序列与上下文数据,提升生成方法可比性
合成数据质量评估仍是保障数据驱动研究中隐私与实用性的重要挑战。本文提出一种评估框架,量化合成数据在还原原始分布特性的同时确保隐私保护的能力。该方法采用留出式基准测试策略,通过低维与高维分布对比、基于嵌入的相似性度量以及最近邻距离指标,实现定量评估。框架支持多种数据类型与结构,包括序列和上下文信息,并通过一组标准化指标提供可解释的质量诊断。这些贡献旨在推动合成数据生成技术基准测试的可复现性与方法一致性。代码已开源:https://github.com/mostly-ai/mostlyai-qa。
原文摘要 · Abstract (English)
Evaluating the quality of synthetic data remains a key challenge for ensuring privacy and utility in data-driven research. In this work, we present an evaluation framework that quantifies how well synthetic data replicates original distributional properties while ensuring privacy. The proposed approach employs a holdout-based benchmarking strategy that facilitates quantitative assessment through low- and high-dimensional distribution comparisons, embedding-based similarity measures, and nearest-neighbor distance metrics. The framework supports various data types and structures, including sequential and contextual information, and enables interpretable quality diagnostics through a set of standardized metrics. These contributions aim to support reproducibility and methodological consistency in benchmarking of synthetic data generation techniques. The code of the framework is available at https://github.com/mostly-ai/mostlyai-qa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。