让可穿戴设备更省电:智能选择何时采样、何时推理
Sense Less, Infer More: Agentic Multimodal Transformers for Edge Medical Intelligence
- 动态控制传感器开关,根据信心和任务重要性自动选通
- 减少48.8%采样次数,同时提升1.9%诊断准确率
- 适合资源受限的可穿戴医疗设备,尤其关注续航与精度平衡
基于边缘的多模态医疗监测需在诊断精度与严苛能耗间取得平衡。持续采集心电图(ECG)、光电容积脉搏波(PPG)、肌电图(EMG)和惯性测量单元(IMU)数据会快速耗尽可穿戴设备电池,通常仅支持不足10小时运行,而现有系统忽视了生理信号中普遍存在的高时间冗余。本文提出自适应多模态智能(AMI),一个端到端框架,联合学习何时采样、如何推理。AMI包含三个组件:(1) 轻量级代理模态控制器,使用可微分的Gumbel-Sigmoid门控机制,根据模型置信度和任务相关性动态选择活跃传感器;(2) 学习型Sigma-Delta感知模块,采用可学习阈值的块级Delta-Sigma操作,跳过时间上冗余的数据样本;(3) 基于单模态基础编码器和跨模态变压器的预测模型,具备时间上下文信息,可在传感器被屏蔽或缺失输入时仍保持鲁棒融合。三者通过多目标损失联合训练,包含分类准确率、稀疏正则化、跨模态对齐和预测编码。AMI支持硬件感知设计,实现动态计算图与掩码操作,带来真实能耗与延迟节省。在MHEALTH、HMC Sleep、WESAD数据集上,传感器使用量降低48.8%,平均准确率优于当前最佳水平1.9%。
原文摘要 · Abstract (English)
Edge-based multimodal medical monitoring requires models that balance diagnostic accuracy with severe energy constraints. Continuous acquisition of ECG, PPG, EMG, and IMU streams rapidly drains wearable batteries, often limiting operation to under 10 hours, while existing systems overlook the high temporal redundancy present in physiological signals. We introduce Adaptive Multimodal Intelligence (AMI), an end-to-end framework that jointly learns when to sense and how to infer. AMI integrates three components: (1) a lightweight Agentic Modality Controller that uses differentiable Gumbel-Sigmoid gating to dynamically select active sensors based on model confidence and task relevance; (2) a Learned Sigma-Delta Sensing module that applies patch-wise Delta-Sigma operations with learnable thresholds to skip temporally redundant samples; and (3) a Foundation-backed Multimodal Prediction Model built on unimodal foundation encoders and a cross-modal transformer with temporal context, enabling robust fusion even under gated or missing inputs. These components are trained jointly via a multi-objective loss combining classification accuracy, sparsity regularization, cross-modal alignment, and predictive coding. AMI is hardware-aware, supporting dynamic computation graphs and masked operations, leading to real energy and latency savings. Across MHEALTH, HMC Sleep, and WESAD datasets, it reduces sensor usage by 48.8% while improving state-of-the-art accuracy by 1.9% on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。