arXiv:2505.05019cs.LGcs.AI2025-05被引 2

优化超参数与领域约束,提升临床试验合成数据质量

Generating Reliable Synthetic Clinical Trial Data: The Role of Hyperparameter Optimization and Domain Constraints

  • 对比九种生成模型,用多指标优化提升数据保真度
  • 复合指标优化使数据更通用,但单靠调参无法避免生存约束违规
  • 预处理和后处理对防止无效数据至关重要,可减少61%的错误

合成临床试验数据为缓解医疗研究中的隐私担忧和数据获取难题提供了新路径。然而,确保合成数据具备高保真度、可用性并遵守领域特定约束仍是关键挑战。本研究系统评估了四种超参数优化(HPO)目标在九种生成模型中的表现,比较了单指标与复合指标优化的效果。结果表明,HPO能持续提升合成数据质量,其中Tab DDPM获益最大,其次为TVAE(60%)、CTGAN(39%)和CTAB-GAN+(38%)。复合指标优化优于单指标,生成更具泛化性的数据。尽管整体质量提高,仅依赖HPO仍无法避免关键临床生存约束的违反。预处理与后处理在降低违规率方面起决定性作用,缺乏稳健处理步骤的模型在61%情况下生成了无效数据。研究强调需结合显式领域知识与HPO以生成高质量合成数据,并为改进合成数据生成提供可操作建议,未来工作需优化指标选择并在更大数据集上验证。

原文摘要 · Abstract (English)

The generation of synthetic clinical trial data offers a promising approach to mitigating privacy concerns and data accessibility limitations in medical research. However, ensuring that synthetic datasets maintain high fidelity, utility, and adherence to domain-specific constraints remains a key challenge. While hyperparameter optimization (HPO) improves generative model performance, the effectiveness of different optimization strategies for synthetic clinical data remains unclear. This study systematically evaluates four HPO objectives across nine generative models, comparing single-metric to compound metric optimization. Our results demonstrate that HPO consistently improves synthetic data quality, with Tab DDPM achieving the largest relative gains, followed by TVAE (60%), CTGAN (39%), and CTAB-GAN+ (38%). Compound metric optimization outperformed single-metric objectives, producing more generalizable synthetic datasets. Despite improving overall quality, HPO alone fails to prevent violations of essential clinical survival constraints. Preprocessing and postprocessing played a crucial role in reducing these violations, as models lacking robust processing steps produced invalid data in up to 61% of cases. These findings underscore the necessity of integrating explicit domain knowledge alongside HPO to generate high-quality synthetic datasets. Our study provides actionable recommendations for improving synthetic data generation, with future work needed to refine metric selection and validate findings on larger datasets.

合成数据临床试验超参数优化领域约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。