用平滑曲线替代复杂训练轨迹,提升医学数据压缩效果。
Improving Clinical Dataset Condensation with Mode Connectivity-based Trajectory Surrogates
- 用二次贝塞尔曲线拟合真实模型训练路径起点与终点
- 在5个临床数据集上表现优于现有方法,模型性能更优
- 适合医疗AI研究者用于高效构建隐私保护合成数据
数据压缩(DC)可生成紧凑且保护隐私的合成临床数据,支持下游临床模型开发。现有方法通过对齐真实数据与合成数据训练过程中的完整随机梯度下降(SGD)轨迹来监督合成数据,但这些轨迹噪声大、曲率高、存储开销大,导致梯度不稳定、收敛慢、内存占用高。本文提出用平滑、低损失的参数化代理替代完整轨迹,采用连接真实训练路径起始与终止状态的二次贝塞尔曲线作为模式连通路径。该方法提供无噪声、低曲率的监督信号,稳定梯度、加速收敛,并消除密集轨迹存储需求。理论证明贝塞尔-模式连接可有效替代SGD路径,实验证明该方法在五个临床数据集上均超越当前最优方案,所生成的压缩数据能支持临床有效模型训练。
原文摘要 · Abstract (English)
Dataset condensation (DC) enables the creation of compact, privacy-preserving synthetic datasets that can match the utility of real patient records, supporting democratised access to highly regulated clinical data for developing downstream clinical models. State-of-the-art DC methods supervise synthetic data by aligning the training dynamics of models trained on real and those trained on synthetic data, typically using full stochastic gradient descent (SGD) trajectories as alignment targets; however, these trajectories are often noisy, high-curvature, and storage-intensive, leading to unstable gradients, slow convergence, and substantial memory overhead. We address these limitations by replacing full SGD trajectories with smooth, low-loss parametric surrogates, specifically quadratic Bézier curves that connect the initial and final model states from real training trajectories. These mode-connected paths provide noise-free, low-curvature supervision signals that stabilise gradients, accelerate convergence, and eliminate the need for dense trajectory storage. We theoretically justify Bézier-mode connections as effective surrogates for SGD paths and empirically show that the proposed method outperforms state-of-the-art condensation approaches across five clinical datasets, yielding condensed datasets that enable clinically effective model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。