arXiv:2604.16955cs.CVcs.AI2026-04

当成像差异大于病情进展时,简单回归模型比复杂生成模型更有效。

Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction

  • 通过输入对齐确保训练与推理一致,提升预测性能。
  • 在视网膜图像上,确定性模型比随机模型表现更好(SSIM提升0.086)。
  • 适合临床慢病预测场景,尤其成像波动大的情况。

从纵向影像预测疾病进展对临床决策和试验设计具有重要意义。近年来方法趋向增加生成模型复杂度,但其必要性仍不明确。我们提出:生成复杂度应匹配任务条件后验中可预测部分的熵,且所有情况下都需保证训练-推理输入对齐。通过两项轻量级测量——原始图像对的任务熵分析与随机模型的后验集中度分析,帮助从业者评估任务所需建模复杂度,再决定框架选择。我们在荧光素眼底照相(FAF)数据集上对比五种条件配置,共享同一架构与训练集,涵盖标准条件扩散、对齐的随机训练和确定性回归。结果显示,训练-推理对齐带来显著提升(delta-SSIM +0.082,SSIM +0.086,p < 0.001),而对齐框架间的差异无临床意义。在两个FAF平台中,两次就诊间变化主要由时间不变的采集变异性主导,而非疾病进展,导致随机模型后验坍缩为有效点,解释了框架等效性。我们训练了一个确定性时间视网膜U-Net(TRU),在三个厂商、两种模态(两个FAF平台和横断面SLO)共28,899只眼上评估,并在三个独立队列上零样本测试。TRU在delta-SSIM、SSIM和PSNR上达到或超过三项已有基线。

原文摘要 · Abstract (English)

Predicting disease progression from longitudinal imaging is useful for clinical decision making and trial design. Recent methods have moved toward increasing generative complexity, but the conditions under which this complexity is necessary remain unclear. We propose that generative complexity should match the entropy of the predictable component of a task's conditional posterior, with training-inference input alignment required in all regimes. Two model-light measurements, a task-entropy analysis on raw image pairs and a posterior-concentration analysis on a stochastic model, let practitioners assess the complexity a task warrants before committing to a modeling framework. We validated this framework on a fundus autofluorescence (FAF) dataset by contrasting five conditioning configurations, sharing one architecture and training set, spanning standard conditional diffusion, inference-aligned stochastic training, and deterministic regression. Training-inference alignment produced large gains (delta-SSIM +0.082, SSIM +0.086, both p < 0.001), while the choice among aligned frameworks produced no clinically meaningful difference across evaluated metrics. Across two FAF platforms, inter-visit change was dominated by time-invariant acquisition variability rather than disease progression, and the stochastic models' posteriors collapsed to an effective point, explaining the framework equivalence. We trained a deterministic Temporal Retinal U-Net (TRU) and evaluated it on 28,899 eyes across three manufacturers and two modalities (two FAF platforms and en-face SLO), with three independent cohorts evaluated zero-shot. TRU matched or exceeded three published baselines on delta-SSIM, SSIM, and PSNR. These findings show that when disease progression is slow compared with acquisition variability, a deterministic regression model matches or outperforms more complex stochastic alternatives.

视网膜图像预测建模确定性模型输入对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。