arXiv:2508.00963cs.LGcs.AI2025-08被引 8

并非模态越多越好,关键在于特征域的互补性。

Rethinking Multimodality: Optimizing Multimodal Deep Learning for Biomedical Signal Classification

  • 通过分析时间、时频、频率域的互补性,优化多模态融合策略。
  • 时频+时间域融合(Hybrid 1)显著提升心电图分类性能,而加入频率域无益。
  • 提出信息论驱动的互补性评估框架,指导高效多模态设计。

本研究重新审视多模态深度学习在生物医学信号分类中的应用,系统分析不同特征域的互补性对模型性能的影响。尽管融合多个模态常被假设可提升准确率,但本文发现并非所有融合都有效。设计并评估了五种模型:三种单模态(1D-CNN用于时间域,2D-CNN用于时频域,1D-CNN-Transformer用于频率域)和两种多模态(Hybrid 1融合1D-CNN与2D-CNN;Hybrid 2整合1D-CNN、2D-CNN与Transformer)。在心电图分类任务中,自助法与贝叶斯推断显示,Hybrid 1在所有指标上均显著优于2D-CNN基线(p<0.05,贝叶斯概率>0.90),证实时间与时频域具有协同互补性。而Hybrid 2引入频率域后未带来提升,甚至出现轻微下降,表明存在表征冗余;该现象经针对性消融实验进一步验证。研究提出‘多模态心电信号深度学习中的互补特征域’理论框架,建立可量化的信息论评估方法,证明最优性能源于融合域间的内在互补性,而非模态数量。

原文摘要 · Abstract (English)

This study proposes a novel perspective on multimodal deep learning for biomedical signal classification, systematically analyzing how complementary feature domains impact model performance. While fusing multiple domains often presumes enhanced accuracy, this work demonstrates that adding modalities can yield diminishing returns, as not all fusions are inherently advantageous. To validate this, five deep learning models were designed, developed, and rigorously evaluated: three unimodal (1D-CNN for time, 2D-CNN for time-frequency, and 1D-CNN-Transformer for frequency) and two multimodal (Hybrid 1, which fuses 1D-CNN and 2D-CNN; Hybrid 2, which combines 1D-CNN, 2D-CNN, and a Transformer). For ECG classification, bootstrapping and Bayesian inference revealed that Hybrid 1 consistently outperformed the 2D-CNN baseline across all metrics (p-values < 0.05, Bayesian probabilities > 0.90), confirming the synergistic complementarity of the time and time-frequency domains. Conversely, Hybrid 2's inclusion of the frequency domain offered no further improvement and sometimes a marginal decline, indicating representational redundancy; a phenomenon further substantiated by a targeted ablation study. This research redefines a fundamental principle of multimodal design in biomedical signal analysis. We demonstrate that optimal domain fusion isn't about the number of modalities, but the quality of their inherent complementarity. This paradigm-shifting concept moves beyond purely heuristic feature selection. Our novel theoretical contribution, "Complementary Feature Domains in Multimodal ECG Deep Learning," presents a mathematically quantifiable framework for identifying ideal domain combinations, demonstrating that optimal multimodal performance arises from the intrinsic information-theoretic complementarity among fused domains.

多模态学习心电图分类特征互补深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。