用耦合模型提升双眼眼底图诊断高度近视的准确率
Copula-enhanced Vision Transformer for high myopia diagnosis through OU UWF fundus images
- 在视觉变换器中加入残差适配器捕捉双眼图像相似与差异
- 提出四维耦合损失和快速蒙特卡洛算法,稳定估计多任务依赖关系
- 特别适合需要同时判断病状和预测眼轴长度的医学影像分析
人工智能辅助的高度近视筛查需联合诊断双眼(OU)高度近视状态及眼轴长度(AL)预测。这一临床需求构成复杂混合型(二分类-连续变量)多任务学习问题,涉及双域(双眼)图像协变量,带来两大挑战:其一,在前沿基础模型中捕捉双眼图像间的非对称性;其二,在给定图像协变量下建模并估计混合型多变量响应间的条件依赖结构。为此,我们提出:其一,在视觉变换器基础模型上引入残差适配器,同步捕捉双眼的相似性与异质性;其二,基于高斯耦合似然的潜变量表达,设计可于PyTorch实现的四维耦合损失,并提出一种计算高效的快速蒙特卡洛期望最大化(fMCEM)算法以估计耦合参数。我们进一步揭示多任务学习中一种称为‘更强协方差现象’的过拟合问题,说明其对耦合参数估计的干扰,并理论证明所提fMCEM算法对此扰动具有数值稳定性。在自建标注的双眼超广角眼底图像数据集及合成数据上的模拟实验表明,该方法在分类与回归任务上均显著且稳定提升预测能力。
原文摘要 · Abstract (English)
The advancement of AI-assisted myopia screening necessitates the joint diagnosis of both-eye (OU) high myopia (HM) status and the prediction of axial length (AL). This clinical requirement introduces a complex mixed-type (binary-continuous) multitask learning task with bi-domain (OU) image covariates, giving rise to two key challenges: i) capture the inter-ocular asymmetry of OU images within a cutting-edge foundation model; ii) model and estimate the conditional dependence structure among mixed-type multivariate responses given image covariates. We address the challenges by: i) imposing residual adapters on the Vision Transformer foundation model to capture the OU similarity and heterogeneity simultaneously; ii) developing a four-dimensional copula loss that is implementable in PyTorch based on a latent variable expression for the Gaussian copula likelihood, and proposing a computationally efficient fast Monte Carlo Expectation Maximization (fMCEM) algorithm to estimate copula parameters. We further formulate a specific overfitting problem called stronger covariance phenomenon in multitask learning. We reveal the disturbance of the phenomenon to estimation of copula parameters and theoretically demonstrate the numerical stability of the proposed fMCEM algorithm against the disturbance. The application to our annotated OU ultra-widefield fundus image dataset and simulation on synthetic data demonstrate that our method stably enhances the predictive capabilities on both classification and regression tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。