arXiv:2507.22274cs.CVcs.LG2025-07被引 19

将手工特征与深度学习结合,提升眼底图像疾病分类准确率

HOG-CNN: Integrating Histogram of Oriented Gradients with Convolutional Neural Networks for Retinal Image Classification

  • 用HOG纹理特征融合CNN深层特征,兼顾局部与语义信息
  • 在多个数据集上表现优异,二分类糖尿病视网膜病变达98.5%准确率
  • 模型轻量可解释,适合资源有限的临床环境部署

眼底图像分析对早期发现糖尿病视网膜病变(DR)、青光眼和年龄相关性黄斑变性(AMD)等眼病至关重要。传统诊断依赖人工判读,耗时且资源密集。为此,我们提出一种基于混合特征提取模型HOG-CNN的自动化、可解释的临床决策支持框架。核心在于融合手工设计的梯度方向直方图(HOG)特征与深度卷积神经网络(CNN)表征,使模型同时捕捉局部纹理模式与高层语义特征。我们在三个公开基准数据集上评估:APTOS 2019(二分类及五分类DR)、ORIGA(青光眼检测)和IC-AMD(AMD诊断)。HOG-CNN表现稳定,二分类DR准确率达98.5%,AUC为99.2;五分类DR的AUC达94.2。在IC-AMD数据集上,准确率为92.8%,精确率为94.8%,AUC为94.5,优于多个前沿模型。在ORIGA上青光眼检测准确率为83.9%,AUC为87.2,尽管数据受限仍具竞争力。附录研究证实了HOG与CNN特征的互补优势。模型轻量且可解释,适用于资源受限的临床场景。结果表明,HOG-CNN是自动眼底疾病筛查中稳健且可扩展的工具。

原文摘要 · Abstract (English)

The analysis of fundus images is critical for the early detection and diagnosis of retinal diseases such as Diabetic Retinopathy (DR), Glaucoma, and Age-related Macular Degeneration (AMD). Traditional diagnostic workflows, however, often depend on manual interpretation and are both time- and resource-intensive. To address these limitations, we propose an automated and interpretable clinical decision support framework based on a hybrid feature extraction model called HOG-CNN. Our key contribution lies in the integration of handcrafted Histogram of Oriented Gradients (HOG) features with deep convolutional neural network (CNN) representations. This fusion enables our model to capture both local texture patterns and high-level semantic features from retinal fundus images. We evaluated our model on three public benchmark datasets: APTOS 2019 (for binary and multiclass DR classification), ORIGA (for Glaucoma detection), and IC-AMD (for AMD diagnosis); HOG-CNN demonstrates consistently high performance. It achieves 98.5\% accuracy and 99.2 AUC for binary DR classification, and 94.2 AUC for five-class DR classification. On the IC-AMD dataset, it attains 92.8\% accuracy, 94.8\% precision, and 94.5 AUC, outperforming several state-of-the-art models. For Glaucoma detection on ORIGA, our model achieves 83.9\% accuracy and 87.2 AUC, showing competitive performance despite dataset limitations. We show, through comprehensive appendix studies, the complementary strength of combining HOG and CNN features. The model's lightweight and interpretable design makes it particularly suitable for deployment in resource-constrained clinical environments. These results position HOG-CNN as a robust and scalable tool for automated retinal disease screening.

医学图像特征融合眼底病可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。