arXiv:2511.14119cs.MMcs.AI2025-11

用手机和智能眼镜实时分析急救现场音视频,提前推荐救治方案。

Real-Time Mobile Video Analytics for Pre-arrival Emergency Medical Services

  • 融合音视频输入,用专用大模型提取症状并估算心率。
  • 多模态联合分析比单一模态更可靠,文本+生命体征预测准确率更高。
  • 适合急救调度、一线医护人员及智慧医疗系统研发者参考。

及时准确的院前视频流与分析对急救服务至关重要。现有急救系统受限于单向视频流和弱分析能力,导致接线员和急救人员在高压环境下需手动处理大量嘈杂冗余信息。我们提出TeleEMS,一种移动端实时视频分析系统,通过融合音频与视频,在急救人员抵达前完成多模态推理。TeleEMS包含客户端与边缘部署的服务器:客户端支持手机、智能眼镜与桌面设备,服务于目击者、途中急救员与911接线员;服务器基于EMS-Stream通信框架实现多方流畅视频传输,并集成三项实时分析模块:(1)基于EMSLlama(领域专用大模型)的语音转症状分析,实现鲁棒症状提取与归一化;(2)采用先进rPPG方法进行视频心率估计;(3)通过PreNet多任务模型实现文本与生命体征联合分析,预测急救流程、用药类型与剂量、操作步骤。评估显示,EMSLlama在准确匹配率上优于GPT-4o(0.89 vs. 0.57),且文本-体征融合显著提升推理鲁棒性,支持可靠的院前干预建议。TeleEMS展现了移动实时视频分析在革新急救运作中的潜力,连接目击者、调度员与急救员,推动下一代智能急救基础设施发展。

原文摘要 · Abstract (English)

Timely and accurate pre-arrival video streaming and analytics are critical for emergency medical services (EMS) to deliver life-saving interventions. Yet, current-generation EMS infrastructure remains constrained by one-to-one video streaming and limited analytics capabilities, leaving dispatchers and EMTs to manually interpret overwhelming, often noisy or redundant information in high-stress environments. We present TeleEMS, a mobile live video analytics system that enables pre-arrival multimodal inference by fusing audio and video into a unified decision-making pipeline before EMTs arrive on scene. TeleEMS comprises two key components: TeleEMS Client and TeleEMS Server. The TeleEMS Client runs across phones, smart glasses, and desktops to support bystanders, EMTs en route, and 911 dispatchers. The TeleEMS Server, deployed at the edge, integrates EMS-Stream, a communication backbone that enables smooth multi-party video streaming. On top of EMSStream, the server hosts three real-time analytics modules: (1) audio-to-symptom analytics via EMSLlama, a domain-specialized LLM for robust symptom extraction and normalization; (2) video-to-vital analytics using state-of-the-art rPPG methods for heart rate estimation; and (3) joint text-vital analytics via PreNet, a multimodal multitask model predicting EMS protocols, medication types, medication quantities, and procedures. Evaluation shows that EMSLlama outperforms GPT-4o (exact-match 0.89 vs. 0.57) and that text-vital fusion improves inference robustness, enabling reliable pre-arrival intervention recommendations. TeleEMS demonstrates the potential of mobile live video analytics to transform EMS operations, bridging the gap between bystanders, dispatchers, and EMTs, and paving the way for next-generation intelligent EMS infrastructure.

急救系统多模态分析实时推理边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。