arXiv:2602.23447eess.IVcs.AI2026-02

用频域扩散生成可控肺部结节图像,提升罕见病检测精度

SALIENT: Frequency-Aware Paired Diffusion for Controllable Long-Tail CT Detection

  • 在小波域进行结构化扩散,分离亮度与细节特征
  • 生成真实感更强的病灶-掩码配对数据,MS-SSIM升至0.83,FID降至46.5
  • 适用于标注少的长尾场景,适合医学影像数据增强研究者

全身影像中罕见病灶检测受极端类别不平衡和低目标体积比制约,导致精确率崩溃,尽管AUROC较高。基于扩散模型的合成增强有潜力,但像素空间扩散计算成本高,现有掩码条件方法缺乏可控属性调节和配对监督。本文提出SALIENT,一种掩码条件的小波域扩散框架,在长尾条件下实现可控的CT增强。不直接在像素空间去噪,而是对离散小波系数进行结构化扩散,显式分离低频亮度与高频结构细节。可学习的频率感知目标解耦目标与背景属性(结构、对比度、边缘保真度),实现可解释且稳定的优化。3D VAE生成多样体数据病灶掩码,半监督教师模型生成切片级伪标签以指导下游检测。SALIENT提升生成真实感,表现为MS-SSIM从0.63升至0.83,FID从118.4降至46.5。下游评估显示,经SALIENT增强训练后,长尾检测性能显著提升,尤其在低预估值和低目标体积比下,AUPRC增益明显。最优合成比例随标注种子量减少从2倍增至4倍,表明在低标签条件下需动态调整增强策略。结果表明,频域感知扩散可在保持计算效率的同时实现可控的精确率恢复。

原文摘要 · Abstract (English)

Detection of rare lesions in whole-body CT is fundamentally limited by extreme class imbalance and low target-to-volume ratios, producing precision collapse despite high AUROC. Synthetic augmentation with diffusion models offers promise, yet pixel-space diffusion is computationally expensive, and existing mask-conditioned approaches lack controllable attribute-level regulation and paired supervision for accountable training. We introduce SALIENT, a mask-conditioned wavelet-domain diffusion framework that synthesizes paired lesion-masking volumes for controllable CT augmentation under long-tail regimes. Instead of denoising in pixel space, SALIENT performs structured diffusion over discrete wavelet coefficients, explicitly separating low-frequency brightness from high-frequency structural detail. Learnable frequency-aware objectives disentangle target and background attributes (structure, contrast, edge fidelity), enabling interpretable and stable optimization. A 3D VAE generates diverse volumetric lesion masks, and a semi-supervised teacher produces paired slice-level pseudo-labels for downstream mask-guided detection. SALIENT improves generative realism, as reflected by higher MS-SSIM (0.63 to 0.83) and lower FID (118.4 to 46.5). In a separate downstream evaluation, SALIENT-augmented training improves long-tail detection performance, yielding disproportionate AUPRC gains across low prevalences and target-to-volume ratios. Optimal synthetic ratios shift from 2x to 4x as labeled seed size decreases, indicating a seed-dependent augmentation regime under low-label conditions. SALIENT demonstrates that frequency-aware diffusion enables controllable, computationally efficient precision rescue in long-tail CT detection.

医学影像扩散模型数据增强长尾检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。