揭示乳腺影像弱监督模型中细粒度特征退化的机制
Gradient-Based Latent Decomposition Reveals Mechanisms of Feature Degradation in Weakly Supervised Mammography
- 用梯度正交分解解析潜空间,分离粗粒度与残差特征
- 仅4.4%潜变量响应监督信号,95.6%依赖脆弱残差空间
- 解释为何病理判别稳定性远低于重建质量,适合医学AI研究者
弱监督分层模型存在持续性不对称:粗粒度病变类型特征在重构中得以保留,而细粒度恶性程度线索却严重退化,直接影响乳腺癌筛查的临床可靠性。本文为分层变分自编码器(H-VAEs)引入基于梯度的正交潜空间分解,将潜空间划分为任务对齐分量(z₁)与正交残差(z_res)。在来自CBIS-DDSM的3,550个乳腺影像兴趣区域上,仅有约4.4%的潜变量幅度与监督梯度对齐,剩余约95.6%分布于残差空间,而细粒度病理预测主要依赖于此。模型在第一阶段达到AUC 0.866,第二阶段为0.552;重构稳定性差距Δ_diag=5%(p=0.005),分类差距Δ_AUC=0.314(p<0.001)。潜变量消融证实两项任务特征均集中于z_res,从结构上解释了为何病理判断稳定性显著低于重构表现。与多实例学习(MIL)和多任务学习(MTL)的对比表明该现象具有跨架构与模态的普遍性。研究揭示,在高维空间中,单一粗粒度监督信号仅激活稀疏的一维潜方向,迫使关键细粒度特征被迫进入易受干扰的残差子空间。
原文摘要 · Abstract (English)
Weakly supervised hierarchical models exhibit a persistent asymmetry: coarse lesion-type features are preserved under reconstruction while fine-grained malignancy cues degrade---a pattern with direct consequences for the clinical reliability of breast cancer screening pipelines. We introduce gradient-based orthogonal latent decomposition for hierarchical Variational Autoencoders~(H-VAEs) to mechanistically explain this asymmetry. The latent space is partitioned into a task-aligned component~($z_1$), shaped by coarse supervisory gradients, and an orthogonal residual~($z_{\text{res}}$) capturing remaining representational capacity. On~3,550 mammographic Regions of Interest~(ROIs) from CBIS-DDSM, only~$\sim$4.4\% of latent magnitude aligns with supervisory gradients, leaving~$\sim$95.6\% in the orthogonal residual upon which fine-grained pathology prediction primarily depends. The model achieves Stage-1~AUC~0.866 and Stage 2~AUC~0.552, with a reconstruction stability gap of $Δ_{\text{diag}}=5\%$ ($p=0.005$) and a classification gap of $Δ_{\text{AUC}}=0.314$ ($p{<}0.001$). Latent ablation confirms that features for both tasks reside heavily in~$z_{\text{res}}$, structurally explaining why reconstruction degrades pathology stability disproportionately. Comparisons with Multi-Instance Learning~(MIL) and Multi-Task Learning~(MTL) confirm generalization across architectures and modalities. These findings reveal that in high-dimensional spaces, a single coarse supervisory signal isolates only a sparse 1D latent direction, forcing critical fine-grained features into the vulnerable residual subspace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。