通过消融分析揭示多模态神经信号融合中跨模态交互的关键作用
IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification
- 构建隔离网络,分步关闭单模态与跨模态注意力路径
- 跨模态交互贡献约30%决策信号,优于线性融合10%以上
- 适用于假肢与神经机器人系统传感器布局设计
手部动作是人体运动系统的主要输出,但其神经肌肉特征的解码仍是基础神经科学和辅助技术(如假肢)的瓶颈。传统人机接口依赖单一生物信号模态,而多模态融合可利用传感器间的互补信息。本文系统比较了线性与基于注意力的融合策略在三种架构中的表现:多模态MLP、多模态Transformer与层级Transformer,评估其在单模态与多模态输入下的性能。实验使用两个公开数据集:NinaPro DB2(sEMG与加速度计)和HD-sEMG 65-Gesture(高密度sEMG与力信号)。在两个数据集上,采用注意力融合的层级Transformer始终表现最佳,在NinaPro DB2上比最优单模态线性融合MLP提升超10%,在HD-sEMG上提升3.7%。为探究模态间交互机制,引入隔离网络,有选择地屏蔽单模态或跨模态注意力路径,量化各类令牌交互对下游决策的贡献。消融实验显示,跨模态交互在各层变压器中贡献约30%的决策信号,凸显注意力驱动融合在挖掘互补信息中的关键作用。研究揭示了多模态融合提升生物信号分类的适用条件,并提供了人类肌肉活动的机制洞察,对神经机器人系统传感器阵列设计具有指导意义。
原文摘要 · Abstract (English)
Hand gestures are a primary output of the human motor system, yet the decoding of their neuromuscular signatures remains a bottleneck for basic neuroscience and assistive technologies such as prosthetics. Traditional human-machine interface pipelines rely on a single biosignal modality, but multimodal fusion can exploit complementary information from sensors. We systematically compare linear and attention-based fusion strategies across three architectures: a Multimodal MLP, a Multimodal Transformer, and a Hierarchical Transformer, evaluating performance on scenarios with unimodal and multimodal inputs. Experiments use two publicly available datasets: NinaPro DB2 (sEMG and accelerometer) and HD-sEMG 65-Gesture (high-density sEMG and force). Across both datasets, the Hierarchical Transformer with attention-based fusion consistently achieved the highest accuracy, surpassing the multimodal and best single-modality linear-fusion MLP baseline by over 10% on NinaPro DB2 and 3.7% on HD-sEMG. To investigate how modalities interact, we introduce an Isolation Network that selectively silences unimodal or cross-modal attention pathways, quantifying each group of token interactions' contribution to downstream decisions. Ablations reveal that cross-modal interactions contribute approximately 30% of the decision signal across transformer layers, highlighting the importance of attention-driven fusion in harnessing complementary modality information. Together, these findings reveal when and how multimodal fusion would enhance biosignal classification and also provides mechanistic insights of human muscle activities. The study would be beneficial in the design of sensor arrays for neurorobotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。