arXiv:2507.07261cs.LGeess.SP2025-07被引 2

融合雷达与可穿戴传感器,提升进食动作识别准确率与鲁棒性

Robust Multimodal Learning Framework For Intake Gesture Detection Using Contactless Radar and Wearable IMU Sensors

  • 设计跨模态注意力网络,融合可穿戴惯性传感器与非接触式雷达数据
  • 在完整数据下比单模态提升4.3%~5.2%的检测准确率,缺一模态仍增1.3%~2.4%
  • 首个在缺失传感器时仍保持性能的多模态进食动作识别框架,适合健康监测应用

自动食物摄入动作检测对饮食监测至关重要,可实现对进食行为的客观、连续追踪,助力改善健康。腕部惯性测量单元(IMU)已广泛用于该任务并取得良好效果,近期非接触式雷达传感器也展现出潜力。本研究探索通过多模态学习融合可穿戴与非接触传感数据以进一步提升检测性能,并解决多模态学习中因某一模态缺失导致鲁棒性下降的核心挑战。为此,提出一种具备跨模态注意力机制的鲁棒多模态时间卷积网络(MM-TCN-CMA),用于融合IMU与雷达数据,提升动作检测能力,并在模态缺失条件下维持性能。构建了一个新数据集,包含52名参与者共52个餐食会话(3,050次进食动作与797次饮水动作),并公开共享。实验结果表明,所提框架在段级F1分数上分别较单模态雷达和IMU模型提升4.3%与5.2%;在模态缺失场景下,仍分别获得1.3%与2.4%的性能增益。这是首个证明可在多模态融合中有效结合IMU与雷达数据并具备强鲁棒性的进食动作识别框架。

原文摘要 · Abstract (English)

Automated food intake gesture detection plays a vital role in dietary monitoring, enabling objective and continuous tracking of eating behaviors to support better health outcomes. Wrist-worn inertial measurement units (IMUs) have been widely used for this task with promising results. More recently, contactless radar sensors have also shown potential. This study explores whether combining wearable and contactless sensing modalities through multimodal learning can further improve detection performance. We also address a major challenge in multimodal learning: reduced robustness when one modality is missing. To this end, we propose a robust multimodal temporal convolutional network with cross-modal attention (MM-TCN-CMA), designed to integrate IMU and radar data, enhance gesture detection, and maintain performance under missing modality conditions. A new dataset comprising 52 meal sessions (3,050 eating gestures and 797 drinking gestures) from 52 participants is developed and made publicly available. Experimental results show that the proposed framework improves the segmental F1-score by 4.3% and 5.2% over unimodal Radar and IMU models, respectively. Under missing modality scenarios, the framework still achieves gains of 1.3% and 2.4% for missing radar and missing IMU inputs. This is the first study to demonstrate a robust multimodal learning framework that effectively fuses IMU and radar data for food intake gesture detection.

动作识别多模态融合健康监测雷达感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。