智能眼镜通过多模态任务模型提升急救人员决策效率
A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
- 构建首个急救场景多模态多任务模型EMSNet,融合文本、生命体征与图像
- 支持五项关键急救任务,准确率显著优于单一模态模型
- 低延迟推理框架使速度提升1.9至11.7倍,适合现场异步数据输入
急救人员在高压环境下需快速做出生死攸关的决策。本文提出EMSGlass智能眼镜系统,基于EMSNet——首个面向急救服务(EMS)的多模态多任务模型,以及专为急救场景设计的低延迟多模态服务框架EMSServe。EMSNet整合文本、生命体征与场景图像,实现对急救事件的统一实时理解。在真实世界多模态急救数据集上训练,可同时支持最多五项关键急救任务,性能优于现有单模态基线模型。基于PyTorch的EMSServe引入模态感知模型拆分器与特征缓存机制,在异构硬件上实现自适应高效推理,有效应对现场模态到达不同步的问题。相比直接使用PyTorch进行多模态推理,其推理速度提升1.9至11.7倍。六名专业急救人员的用户研究显示,EMSGlass显著提升了实时态势感知、决策速度与操作效率。定性反馈为下一代人工智能赋能的急救系统提供了可行动方向,推动多模态智能与真实应急响应流程深度融合。
原文摘要 · Abstract (English)
Emergency Medical Technicians (EMTs) operate in high-pressure environments, making rapid, life-critical decisions under heavy cognitive and operational loads. We present EMSGlass, a smart-glasses system powered by EMSNet, the first multimodal multitask model for Emergency Medical Services (EMS), and EMSServe, a low-latency multimodal serving framework tailored to EMS scenarios. EMSNet integrates text, vital signs, and scene images to construct a unified real-time understanding of EMS incidents. Trained on real-world multimodal EMS datasets, EMSNet simultaneously supports up to five critical EMS tasks with superior accuracy compared to state-of-the-art unimodal baselines. Built on top of PyTorch, EMSServe introduces a modality-aware model splitter and a feature caching mechanism, achieving adaptive and efficient inference across heterogeneous hardware while addressing the challenge of asynchronous modality arrival in the field. By optimizing multimodal inference execution in EMS scenarios, EMSServe achieves 1.9x -- 11.7x speedup over direct PyTorch multimodal inference. A user study evaluation with six professional EMTs demonstrates that EMSGlass enhances real-time situational awareness, decision-making speed, and operational efficiency through intuitive on-glass interaction. In addition, qualitative insights from the user study provide actionable directions for extending EMSGlass toward next-generation AI-enabled EMS systems, bridging multimodal intelligence with real-world emergency response workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。