arXiv:2604.15961cs.LG2026-04

提出一套评估合成健康数据质量的方法,解决缺乏统一标准的问题。

Evaluating quality in synthetic data generation for large tabular health datasets

论文配图:Evaluating quality in synthetic data generation for large tabular health datasets
图 1 · 摘自论文原文
  • 用统一调参流程对比7种主流模型在4个不同规模数据集上的表现
  • 设计可视化对齐的联合分布保真度评估法,可跨数据集通用
  • 针对德国癌症登记数据揭示模型在医学领域约束下的局限性

目前学术界在合成数据的质量评估方面尚无统一标准,尤其缺乏对大型健康数据集(如历史流行病学数据)的基准测试。本研究评估了来自主要机器学习家族的七种近期模型,使用四个具有不同规模的数据集进行实验。为确保公平比较,我们为每种模型在每个数据集上系统地进行了超参数调优。本文提出一种评估合成数据联合分布保真度的方法,将指标与单一图表可视化对齐,适用于任意数据集,并辅以针对德国癌症登记处流行病学数据的领域特定分析。分析揭示了模型在严格遵守医学领域约束时面临的挑战。希望该方法能成为指导合成器选择的基础框架,并对所有参与发布合成数据的人员保持可及性。

原文摘要 · Abstract (English)

There is no consensus in the field of synthetic data on concise metrics for quality evaluations or benchmarks on large health datasets, such as historical epidemiological data. This study presents an evaluation of seven recent models from major machine learning families. The models were evaluated using four different datasets, each with a distinct scale. To ensure a fair comparison, we systematically tuned the hyperparameters of each model for each dataset. We propose a methodology for evaluating the fidelity of synthesized joint distributions, aligning metrics with visualization on a single plot. This method is applicable to any dataset and is complemented by a domain-specific analysis of the German Cancer Registries' epidemiological dataset. The analysis reveals the challenges models face in strictly adhering to the medical domain. We hope this approach will serve as a foundational framework for guiding the selection of synthesizers and remain accessible to all stakeholders involved in releasing synthetic datasets.

合成数据健康数据评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。