通过渐进式轨迹匹配,高效压缩医学图像数据集。
High-Order Progressive Trajectory Matching for Medical Image Dataset Distillation
- 基于参数轨迹的几何结构设计形状势能函数
- 采用由简到繁的匹配策略提升蒸馏效果
- 在保护隐私前提下保持接近原始数据的模型精度
医学图像分析因隐私法规和机构流程复杂,面临数据共享难题。数据蒸馏提供了解决方案,通过合成紧凑数据集来保留真实大型医学数据集的核心信息。轨迹匹配已成为数据蒸馏的有前景方法,但现有方法多关注终端状态,忽视优化过程中的中间状态信息。本文提出一种形状相关的势能函数以捕捉参数轨迹的几何结构,并设计由易到难的匹配策略,逐步处理不同复杂度的参数。在医学图像分类任务上的实验表明,该方法在保持隐私的同时显著提升蒸馏性能,模型精度与使用原始数据训练相当。代码已开源:https://github.com/Bian-jh/HoP-TM。
原文摘要 · Abstract (English)
Medical image analysis faces significant challenges in data sharing due to privacy regulations and complex institutional protocols. Dataset distillation offers a solution to address these challenges by synthesizing compact datasets that capture essential information from real, large medical datasets. Trajectory matching has emerged as a promising methodology for dataset distillation; however, existing methods primarily focus on terminal states, overlooking crucial information in intermediate optimization states. We address this limitation by proposing a shape-wise potential that captures the geometric structure of parameter trajectories, and an easy-to-complex matching strategy that progressively addresses parameters based on their complexity. Experiments on medical image classification tasks demonstrate that our method improves distillation performance while preserving privacy and maintaining model accuracy comparable to training on the original datasets. Our code is available at https://github.com/Bian-jh/HoP-TM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。