不靠频率分解,用几何交互实现高效眼底图像多标签诊断
Less is More in Semantic Space: Intrinsic Decoupling via Clifford-M for Fundus Image Classification
- 用克莱夫福德几何乘积构建轻量级双分辨率结构,直接建模跨尺度特征关系
- 仅0.85M参数即达0.8142平均AUC-ROC,超越更大模型
- 无需预训练,在跨数据集测试中仍保持稳定性能,适合资源受限场景
多标签眼底诊断需同时捕捉细微病灶与大范围视网膜结构。现有模型常通过显式频域分解应对,但实验表明此类设计提升有限:用八度卷积替代原双分辨率主干使参数增加35%,计算量增至2.23倍,却未提高平均准确率;基于小波的固定变体表现更差。为此,我们提出Clifford-M,一种轻量骨干网络,以稀疏几何交互取代前馈扩展与频域分裂模块。该模型基于克莱夫福德风格滚动积,以线性复杂度联合建模对齐与结构变化,实现在紧凑双分辨率架构下的高效跨尺度融合与自精炼。无需预训练,Clifford-M在ODIR-5K上取得0.8142的平均AUC-ROC和0.5481的平均宏F1(最优阈值),仅用0.85M参数,显著优于同类中等规模CNN基线。在未微调的RFMiD数据集上,其宏AUC为0.7425±0.0198,微AUC为0.7610±0.0344,展现出良好跨数据集泛化能力。结果表明,只要核心特征交互能直接捕获多尺度结构,即可实现高性能且高效的视网膜图像诊断,无需显式频域工程。
原文摘要 · Abstract (English)
Multi-label fundus diagnosis requires features that capture both fine-grained lesions and large-scale retinal structure. Many multi-scale medical vision models address this challenge through explicit frequency decomposition, but our ablation studies show that such heuristics provide limited benefit in this setting: replacing the proposed simple dual-resolution stem with Octave Convolution increased parameters by 35% and computation by a 2.23-fold increase in computation; without improving mean accuracy, while a fixed wavelet-based variant performed substantially worse. Motivated by these findings, we propose Clifford-M, a lightweight backbone that replaces both feed-forward expansion and frequency-splitting modules with sparse geometric interaction. The model is built on a Clifford-style rolling product that jointly captures alignment and structural variation with linear complexity, enabling efficient cross-scale fusion and self-refinement in a compact dual-resolution architecture. Without pre-training, Clifford-M achieves a mean AUC-ROC of 0.8142 and a mean macro-F1 (optimal threshold) of 0.5481 on ODIR-5K using only 0.85M parameters, outperforming substantially larger mid-scale CNN baselines under the same training protocol. When evaluated on RFMiD without fine-tuning, it attains 0.7425 +/- 0.0198 macro AUC and 0.7610 +/- 0.0344 micro AUC, indicating reasonable robustness to cross-dataset shift. These results suggest that competitive and efficient fundus diagnosis can be achieved without explicit frequency engineering, provided that the core feature interaction is designed to capture multi-scale structure directly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。