arXiv:2512.14491cs.AI2025-12

提出稀疏多模态Transformer,高效分类阿尔茨海默病。

Sparse Multi-Modal Transformer with Masking for Alzheimer's Disease Classification

  • 用聚类稀疏注意力降低计算量,近线性复杂度。
  • 模态掩码提升对缺失数据的鲁棒性,准确率保持稳定。
  • 适合资源受限场景下的可扩展智能系统部署。

基于Transformer的多模态智能系统因密集自注意力导致高计算与能耗,限制了在资源受限环境下的可扩展性。本文提出SMMT稀疏多模态Transformer架构,基于级联多模态框架,引入基于聚类的稀疏注意力机制,实现近线性计算复杂度,并采用模态级掩码策略增强对不完整输入的鲁棒性。以ADNI数据集上的阿尔茨海默病分类为案例研究,实验表明SMMT在保持竞争性预测性能的同时,显著降低训练时间、内存占用与能耗,验证其作为资源感知型架构组件在可扩展智能系统中的适用性。

原文摘要 · Abstract (English)

Transformer-based multi-modal intelligent systems often suffer from high computational and energy costs due to dense self-attention, limiting their scalability under resource constraints. This paper presents SMMT, a sparse multi-modal transformer architecture designed to improve efficiency and robustness. Building upon a cascaded multi-modal transformer framework, SMMT introduces cluster-based sparse attention to achieve near linear computational complexity and modality-wise masking to enhance robustness against incomplete inputs. The architecture is evaluated using Alzheimer's Disease classification on the ADNI dataset as a representative multi-modal case study. Experimental results show that SMMT maintains competitive predictive performance while significantly reducing training time, memory usage, and energy consumption compared to dense attention baselines, demonstrating its suitability as a resource-aware architectural component for scalable intelligent systems.

多模态稀疏注意力阿尔茨海默病高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。