arXiv:2606.00588cs.CV2026-06

基于早期治疗反应预测糖尿病黄斑水肿患者长期视力变化

Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting

论文配图:Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting
图 1 · 摘自论文原文
  • 融合基线与1个月OCT影像及临床指标,建模治疗早期反应
  • 24个月视力预测平均误差仅0.1246,决定系数达0.6064
  • 适合眼科医生做个性化随访规划,支持抗VEGF治疗决策

糖尿病黄斑水肿(DME)患者接受抗VEGF治疗后的长期视力预后对患者沟通、期望管理及随访计划至关重要。但临床上常需仅凭早期治疗数据估算长期视觉轨迹,导致预测困难。现有基于OCT的学习方法多聚焦短期反应或单一终点预测,而利用早期纵向观测数据建模多时间点视力轨迹的研究仍不足。本研究构建了包含188例抗VEGF治疗的DME患者的真实世界队列,包含基线和治疗1个月的OCT扫描、表格化OCT衍生生物标志物及非影像学临床变量。仅使用这些早期数据,我们提出多时域视力预测任务,目标是预测3、6、12、18、24个月的视觉敏锐度,对应临床有意义的随访间隔。我们提出ReVA框架,整合基线与1个月OCT的结构特征及表格变量,捕捉疾病基线状态与早期治疗反应。ReVA采用空间注意力保留局部预后影像特征,并通过依赖感知表格编码器建模临床变量间交互关系。多模态表征融合后预测个体化长期视力轨迹。该框架在24个月视力预测中达到MAE=0.1246,RMSE=0.1621,R²=0.6064,且各预测时域表现稳定。结果表明,纳入早期治疗反应信号可实现临床意义的长期视力预测,支持抗VEGF治疗的自动化决策支持。

原文摘要 · Abstract (English)

Long-term visual acuity (VA) outcomes after anti-VEGF therapy are central to patient counseling, expectation setting, and follow-up planning in diabetic macular edema (DME). However, in clinical practice, physicians must often estimate long-term visual trajectories based only on early post-treatment findings, making reliable prognostication difficult. Although prior OCT-based learning approaches have largely focused on short-term response or single-endpoint prediction, modeling VA trajectories across multiple future time points from early longitudinal observations remains insufficiently explored. In this study, we assembled a real-world cohort of 188 anti-VEGF--treated DME patients with paired baseline and month-1 OCT scans, along with tabular OCT-derived biomarkers and non-imaging clinical variables. Using only these early data, we formulate a multi-horizon VA forecasting problem aimed at predicting visual outcomes at 3, 6, 12, 18, and 24 months, reflecting clinically meaningful follow-up intervals. We propose \textbf{ReVA}, a response-aware multimodal framework that integrates structural features from baseline and month-1 OCT with the tabular variables to capture baseline disease status and early treatment response. ReVA uses spatial attention to preserve localized prognostic imaging features and a dependency-aware tabular encoder to model interactions among clinical variables. These multimodal representations are fused to predict patient-specific long-term visual acuity trajectories. The proposed framework achieves MAE $=0.1246$, RMSE $=0.1621$, and $R^2=0.6064$ for 24-month VA prediction, with consistent performance across all forecast horizons. Our findings show that incorporating early treatment-response signals enables clinically meaningful long-term visual acuity forecasting, supporting data-driven decision support for routine anti-VEGF management.

视力预测多模态学习OCT分析DME

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。