用眼动数据引导神经网络,提升医疗图像分析的准确性与可解释性
EG-SpikeFormer: Eye-Gaze Guided Transformer on Spiking Neural Networks for Medical Image Analysis

- 基于脉冲神经网络,融合眼动数据指导注意力聚焦
- 在有限数据下显著减少模型偏差,提升诊断可靠性
- 适合临床场景,兼顾能效与可解释性,推动类脑计算落地医疗
类脑计算作为高效节能的人工智能替代方案,主要依赖脉冲神经网络(SNN)在类脑硬件上实现。尽管基于SNN的卷积神经网络和Transformer架构已取得进展,但其在医学影像领域的应用仍较薄弱。本文提出EG-SpikeFormer,一种专为临床任务设计的SNN架构,通过引入眼动数据引导模型关注医学图像中具有诊断意义的区域。该方法有效缓解了传统模型在数据有限时的捷径学习问题,尤其在对模型可靠性、泛化性和透明度要求高的场景下表现突出。EG-SpikeFormer在医学图像预测任务中兼具优异能效与性能,通过多模态信息对齐增强了临床相关性。眼动数据的融入提升了模型可解释性与泛化能力,为类脑计算在医疗领域的应用开辟新路径。
原文摘要 · Abstract (English)
Neuromorphic computing has emerged as a promising energy-efficient alternative to traditional artificial intelligence, predominantly utilizing spiking neural networks (SNNs) implemented on neuromorphic hardware. Significant advancements have been made in SNN-based convolutional neural networks (CNNs) and Transformer architectures. However, neuromorphic computing for the medical imaging domain remains underexplored. In this study, we introduce EG-SpikeFormer, an SNN architecture tailored for clinical tasks that incorporates eye-gaze data to guide the model's attention to the diagnostically relevant regions in medical images. Our developed approach effectively addresses shortcut learning issues commonly observed in conventional models, especially in scenarios with limited clinical data and high demands for model reliability, generalizability, and transparency. Our EG-SpikeFormer not only demonstrates superior energy efficiency and performance in medical image prediction tasks but also enhances clinical relevance through multi-modal information alignment. By incorporating eye-gaze data, the model improves interpretability and generalization, opening new directions for applying neuromorphic computing in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。