arXiv:2502.16914cs.SDcs.AI2025-02中稿 · but not published …被引 6

用CNN和Transformer融合模型提升心跳音诊断准确率

ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on Heart Sounds

  • 采用专家混合框架融合CNN与视觉变压器
  • 在心音分类中达到97.52%准确率,超越单一模型
  • 适合心血管疾病智能诊断研究者参考

本研究探索了视觉变换器(ViT)原理在音频分析中的应用,聚焦于心音识别。本文提出ENACT-Heart——一种基于CNN与ViT互补优势的集成方法,通过专家混合(MoE)框架实现97.52%的分类准确率,显著优于单独使用ViT(93.88%)或CNN(95.45%)的表现,证明了集成方法在心血管健康监测与诊断中的潜力。

原文摘要 · Abstract (English)

This study explores the application of Vision Transformer (ViT) principles in audio analysis, specifically focusing on heart sounds. This paper introduces ENACT-Heart - a novel ensemble approach that leverages the complementary strengths of Convolutional Neural Networks (CNN) and ViT through a Mixture of Experts (MoE) framework, achieving a remarkable classification accuracy of 97.52%. This outperforms the individual contributions of ViT (93.88%) and CNN (95.45%), demonstrating the potential for enhanced diagnostic accuracy in cardiovascular health monitoring. These results demonstrate the potential of ensemble methods in enhancing classification performance for cardiovascular health monitoring and diagnosis.

心音分析Transformer集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。