提出新型视网膜分区模型,让糖尿病视网膜病变诊断更可解释。
HSQ-VLM: A Novel Spatially-Constrained Quadrant Segmentation VLM Model for Explainability in Diabetic Retinopathy

- 基于黄斑中心坐标系动态划分四个象限,实现解剖定位对齐。
- 出血检测敏感度达99.6%,微动脉瘤达96.4%,边界模糊误差显著降低。
- 适合临床医生验证AI决策依据,提升诊疗可信度。
糖尿病视网膜病变(DR)是全球失明的主要原因,但当前诊断AI存在黑箱问题。尽管深度学习模型分类准确率高,却缺乏能精准描述病灶位置与分布的可解释方法。为此,本文提出HSQ-VLM,一种基于视网膜图像的新型象限分割框架,采用地标锚定的笛卡尔交叉注意力机制,将视觉特征提取与结构化临床推理统一。不同于传统任意切分方式,该方法通过四象限拓扑潜空间分割(TLP)动态对齐视网膜特征与以黄斑为中心的坐标系统,使视觉语言模型能生成量化病理、具解剖精度的自然语言报告。在包含3,500张高分辨率眼底图像的数据集上,该方法对出血的检测敏感度达99.6%,对微动脉瘤为96.4%,且显著减少边界模糊错误。
原文摘要 · Abstract (English)
Diabetic Retinopathy (DR) is an aggressive retinal disease and a leading cause of global blindness, yet its clinical management is currently hindered by the black-box nature of diagnostic AI. While deep learning models achieve high classification accuracy, there is a critical lack of explainability methods capable of detailing the exact anatomical landmarks and lesion distributions that lead to a clinical decision for DR. Therefore, we propose HSQ-VLM, a novel quadrant segmentation pipeline on fundus images that utilizes a Landmark-Anchored Cartesian Cross-Attention mechanism to unify visual feature extraction with structured clinical reasoning. Unlike traditional methods that rely on arbitrary image partitioning, our pipeline implements 4-quadrant Topological Latent Partitioning (TLP) to dynamically align retinal features with a fovea-centered coordinate system. This allows the Vision-Language Model to generate natural language reports that quantify pathology with anatomical precision. On a dataset of 3,500 high-resolution fundus images, this innovative methodology achieved a lesion detection sensitivity of 99.6% for hemorrhages and 96.4% for microaneurysms, while demonstrating a significant reduction in boundary-ambiguity errors compared to standard segmentation baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。