arXiv:2505.12887eess.IVcs.CV2025-05被引 5

用文字描述生成高分辨率眼底图像,提升疾病细节表现力。

RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

  • 基于视觉语言模型自动生成140万条带描述的眼底图数据
  • 合成图像62.07%被眼科医生误认为真实图像,提升糖尿病视网膜病变与青光眼检测准确率5%-10%
  • 支持细粒度病灶控制,适合需要精准病理模拟的研究场景

高质量标注眼底影像数据稀缺,制约了眼科机器学习模型的发展。现有彩色眼底照片(CFPs)合成方法多依赖预设疾病标签,难以体现细微解剖差异、早期病变阶段及多样病理特征。为此,我们提出创新流程,构建大规模带描述的眼底图像数据集RetinaLogos-1400k,包含140万条样本。该数据集利用视觉语言模型(VLM)对视盘形态、血管分布、神经纤维层及病灶等关键结构进行描述。在此基础上,我们设计RetinaLogos三步训练框架,实现对眼底图像的细粒度语义控制,准确捕捉疾病进展阶段、微小解剖变异及特定病灶类型。大量实验表明,该方法在多个数据集上表现优异,62.07%的文字驱动合成CFPs被眼科医生误判为真实图像;合成数据可使糖尿病视网膜病变分级和青光眼检测准确率提升5%-10%。代码已开源:https://github.com/uni-medical/retina-text2cfp。

原文摘要 · Abstract (English)

The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour Fundus Photographs (CFPs) largely rely on predefined disease labels, which restricts their ability to generate images that reflect fine-grained anatomical variations, subtle disease stages, and diverse pathological features beyond coarse class categories. To overcome these challenges, we first introduce an innovative pipeline that creates a large-scale, captioned retinal dataset comprising 1.4 million entries, called RetinaLogos-1400k. Specifically, RetinaLogos-1400k uses the visual language model(VLM) to describe retinal conditions and key structures, such as optic disc configuration, vascular distribution, nerve fibre layers, and pathological features. Building on this dataset, we employ a novel three-step training framework, RetinaLogos, which enables fine-grained semantic control over retinal images and accurately captures different stages of disease progression, subtle anatomical variations, and specific lesion types. Through extensive experiments, our method demonstrates superior performance across multiple datasets, with 62.07% of text-driven synthetic CFPs indistinguishable from real ones by ophthalmologists. Moreover, the synthetic data improves accuracy by 5%-10% in diabetic retinopathy grading and glaucoma detection. Codes are available at https://github.com/uni-medical/retina-text2cfp.

眼底图像生成文本到图像医学影像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。