统一扩散模型解决多任务医学图像分割中的标签冲突与梯度失衡问题。
SNR-Adaptive Unified Diffusion for Multi-Task Medical Image Segmentation

- 通过11通道任务专属输出空间避免跨任务梯度反向。
- 在低信噪比阶段抑制领域偏差,提升任务引导精度。
- 适合需要跨数据集、多模态联合训练的医学影像研究者。
临床心脏成像流程目前为每种数据集和模态部署独立模型,导致训练成本重复且难以共享知识。将半监督学习、无监督域适应与域泛化整合到单一模型中成为迫切需求,但直接联合训练存在根本障碍:不同数据集间标签语义冲突使左心室(LA)Dice从90.31%降至83.38%,而复杂度不一的任务间梯度失衡会压制较弱任务。本文提出UniT-Diff统一扩散分割框架,通过三项机制解决上述问题:11通道任务专属输出空间物理划分标签类别,从根本上消除跨任务梯度符号反转;基于信噪比自适应的任务条件(SATC)根据当前扩散步的对数信噪比缩放任务令牌,在粗去噪阶段抑制领域偏差,并在信号清晰时恢复完整任务指导;任务类型感知条件丢弃(TTACD)永久移除域泛化输入的任务令牌,使其通过共享中性路径,依赖跨数据集心脏解剖结构而非源厂商统计特征。在单一参数设置下,UniT-Diff在三个基准测试上均超越独立训练的专用基线:LA提升+0.87%,MMWHS提升+1.77%,MNMS提升+0.88%。
原文摘要 · Abstract (English)
Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and precluding knowledge sharing across anatomically related tasks. Consolidating semi-supervised learning, unsupervised domain adaptation, and domain generalisation into one model is therefore a practical necessity, yet naive joint training exposes a fundamental barrier: conflicting label semantics between datasets collapse LA Dice from 90.31\% to 83.38\%, while gradient imbalance across tasks of unequal complexity suppresses the weaker tasks throughout training. We present UniT-Diff, a unified diffusion segmentation framework that resolves these conflicts through three targeted mechanisms. An 11-channel task-specific output space physically partitions label categories, eliminating cross-task gradient sign reversal by construction. SNR-Adaptive Task Conditioning (SATC) scales the task token by the log signal-to-noise ratio of the current diffusion timestep, suppressing domain-specific bias during coarse denoising and restoring full task guidance as the signal clears. Task-Type-Aware Conditional Dropout (TTACD) permanently removes the task token for domain-generalisation inputs, routing them through a shared neutral pathway that draws on cross-dataset cardiac anatomy rather than source-vendor statistics. Under a single parameter set, UniT-Diff surpasses independently trained task-specific baselines on all three benchmarks simultaneously: +0.87\% on LA, +1.77\% on MMWHS, and +0.88\% on MNMS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。