arXiv:2603.14039cs.CV2026-03

构建可预测眼病进展的动态眼病模型,提升多模态影像分析稳定性。

EyeWorld: A Generative World Model of Ocular State and Dynamics

  • 将眼睛视为动态系统,学习跨模态稳定的潜在眼态表征。
  • 实现多模态影像的精细分割、结构保持转换与质量鲁棒增强。
  • 支持基于时间的病变进展预测,适合临床诊疗与长期随访场景。

眼科决策依赖于多模态影像中细微病灶特征的时序解读,但现有医学基础模型多为静态,易受模态与采集差异影响。本文提出EyeWorld,一种生成式眼病世界模型,将眼睛建模为部分可观测的动力学系统。该模型学习跨模态共享的稳定眼态潜在表示,统一实现细粒度解析、结构保持的跨模态转换与质量鲁棒增强。纵向监督进一步支持时间条件下的状态转移,可预测具有临床意义的病变进展,同时保持解剖结构稳定。通过从静态表征学习转向显式动力学建模,EyeWorld为医学领域提供了统一的鲁棒多模态解读与以预后为导向的仿真方法。

原文摘要 · Abstract (English)

Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade under modality and acquisition shifts. Here we introduce EyeWorld, a generative world model that conceptualizes the eye as a partially observed dynamical system grounded in clinical imaging. EyeWorld learns an observation-stable latent ocular state shared across modalities, unifying fine-grained parsing, structure-preserving cross-modality translation and quality-robust enhancement within a single framework. Longitudinal supervision further enables time-conditioned state transitions, supporting forecasting of clinically meaningful progression while preserving stable anatomy. By moving from static representation learning to explicit dynamical modeling, EyeWorld provides a unified approach to robust multimodal interpretation and prognosis-oriented simulation in medicine.

生成模型眼病诊断动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。