arXiv:2603.09356cs.LGcs.AI2026-03

用合成数据让临床模型共享更安全,不依赖神经网络。

Democratising Clinical AI through Dataset Condensation for Classical Clinical Models

  • 用零阶优化方法让非可微临床模型也能做数据压缩。
  • 六组临床数据验证:合成数据能保持模型性能且隐私保护有效。
  • 适合想共享敏感医疗数据的研究者或机构使用。

数据压缩(DC)通过学习紧凑的合成数据集,使模型在少量数据上训练仍能达到全量数据的效果,注重实用性能而非分布相似性。尽管常用于提升计算效率,DC在医疗领域也有助于数据民主化,尤其结合差分隐私后,合成数据可替代真实病历实现安全共享。然而现有方法依赖可微神经网络,难以适配广泛使用的决策树、Cox回归等不可微临床模型。本文提出一种差分隐私保障的零阶优化框架,仅通过函数评估即可将DC扩展至非可微模型。在六组数据集(涵盖分类与生存分析任务)上的实验表明,该方法生成的压缩数据集既能保持模型性能,又提供有效的差分隐私保护,实现无需暴露真实患者信息的模型无关临床数据共享。

原文摘要 · Abstract (English)

Dataset condensation (DC) learns a compact synthetic dataset that enables models to match the performance of full-data training, prioritising utility over distributional fidelity. While typically explored for computational efficiency, DC also holds promise for healthcare data democratisation, especially when paired with differential privacy, allowing synthetic data to serve as a safe alternative to real records. However, existing DC methods rely on differentiable neural networks, limiting their compatibility with widely used clinical models such as decision trees and Cox regression. We address this gap using a differentially private, zero-order optimisation framework that extends DC to non-differentiable models using only function evaluations. Empirical results across six datasets, including both classification and survival tasks, show that the proposed method produces condensed datasets that preserve model utility while providing effective differential privacy guarantees - enabling model-agnostic data sharing for clinical prediction tasks without exposing sensitive patient information.

临床AI数据压缩差分隐私模型无关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。