提出首个用于医学图像增强的低层级数据蒸馏方法,兼顾隐私与质量。
Low-Level Dataset Distillation for Medical Image Enhancement
- 基于解剖相似性构建共享先验,再通过个性化模块生成患者特异数据。
- 蒸馏数据能还原像素级细节,支持多种低级任务训练对。
- 仅分享抽象数据,保护患者隐私,适合医疗部署场景。
医学图像增强具有临床价值,但现有方法需大规模数据学习复杂的像素级映射,带来高昂的训练与存储成本。尽管数据蒸馏(DD)可缓解负担,但现有方法多针对高阶任务,依赖多样本共标签的映射结构。而低阶任务为多对多的像素级映射,要求精细保真,使低阶数据蒸馏成为欠定问题。为此,本文提出首个用于医学图像增强的低层级数据蒸馏方法。首先利用患者间解剖相似性,以代表性患者构建共享解剖先验,作为各患者蒸馏数据的初始化。随后通过结构保持个性化生成(SPG)模块,将患者特异性解剖信息融入蒸馏数据,同时保留像素级保真度。针对不同低级任务,用蒸馏数据构建特定的高质量与低质量训练对。通过将蒸馏对训练网络所得梯度,与真实患者数据训练梯度对齐,注入患者专属知识。下游用户无法访问原始患者数据,仅接收含抽象训练信息的蒸馏数据,不包含患者特异性细节,有效保障隐私。
原文摘要 · Abstract (English)
Medical image enhancement is clinically valuable, but existing methods require large-scale datasets to learn complex pixel-level mappings. However, the substantial training and storage costs associated with these datasets hinder their practical deployment. While dataset distillation (DD) can alleviate these burdens, existing methods mainly target high-level tasks, where multiple samples share the same label. This many-to-one mapping allows distilled data to capture shared semantics and achieve information compression. In contrast, low-level tasks involve a many-to-many mapping that requires pixel-level fidelity, making low-level DD an underdetermined problem, as a small distilled dataset cannot fully constrain the dense pixel-level mappings. To address this, we propose the first low-level DD method for medical image enhancement. We first leverage anatomical similarities across patients to construct the shared anatomical prior based on a representative patient, which serves as the initialization for the distilled data of different patients. This prior is then personalized for each patient using a Structure-Preserving Personalized Generation (SPG) module, which integrates patient-specific anatomical information into the distilled dataset while preserving pixel-level fidelity. For different low-level tasks, the distilled data is used to construct task-specific high- and low-quality training pairs. Patient-specific knowledge is injected into the distilled data by aligning the gradients computed from networks trained on the distilled pairs with those from the corresponding patient's raw data. Notably, downstream users cannot access raw patient data. Instead, only a distilled dataset containing abstract training information is shared, which excludes patient-specific details and thus preserves privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。