arXiv:2508.09362cs.CVcs.AI2025-08ICCV

用注意力融合多模态数据,提升手语识别准确率

FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition

  • 四路时空网络并行处理视觉与雷达数据,动态融合特征
  • 在意大利手语数据集上达99.44%准确率,优于现有方法
  • 适合医疗场景中复杂手语识别,代码开源可复现

医疗沟通中的手语准确识别面临重大挑战,需能解析复杂多模态手势的框架。为此,我们提出FusionEnsemble-Net,一种基于注意力机制的时空网络集成模型,通过动态融合视觉与运动数据提升识别精度。该方法同步处理RGB视频与距离多普勒雷达模态,经由四条不同路径的时空网络处理,每条路径均采用注意力融合模块持续整合双模态特征,再输入分类器集成。最终,四个融合通道的输出在集成分类头中合并,增强模型鲁棒性。实验表明,FusionEnsemble-Net在大规模MultiMeDaLIS数据集(意大利手语)上达到99.44%测试准确率,显著优于现有方法。结果表明,通过注意力融合统一多样时空网络的集成框架,可有效应对复杂、多模态孤立手势识别任务。源码已公开:https://github.com/rezwanh001/Multimodal-Isolated-Italian-Sign-Language-Recognition。

原文摘要 · Abstract (English)

Accurate recognition of sign language in healthcare communication poses a significant challenge, requiring frameworks that can accurately interpret complex multimodal gestures. To deal with this, we propose FusionEnsemble-Net, a novel attention-based ensemble of spatiotemporal networks that dynamically fuses visual and motion data to enhance recognition accuracy. The proposed approach processes RGB video and range Doppler map radar modalities synchronously through four different spatiotemporal networks. For each network, features from both modalities are continuously fused using an attention-based fusion module before being fed into an ensemble of classifiers. Finally, the outputs of these four different fused channels are combined in an ensemble classification head, thereby enhancing the model's robustness. Experiments demonstrate that FusionEnsemble-Net outperforms state-of-the-art approaches with a test accuracy of 99.44% on the large-scale MultiMeDaLIS dataset for Italian Sign Language. Our findings indicate that an ensemble of diverse spatiotemporal networks, unified by attention-based fusion, yields a robust and accurate framework for complex, multimodal isolated gesture recognition tasks. The source code is available at: https://github.com/rezwanh001/Multimodal-Isolated-Italian-Sign-Language-Recognition.

手语识别多模态融合注意力机制医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。