AI自动分析远程监测数据,精准识别紧急情况,效率远超医生。
From Days to Minutes: An Autonomous AI Agent Achieves Reliable Clinical Triage in Remote Patient Monitoring
- 用多步推理和21种临床工具自动分析生命体征,实现智能分诊。
- 紧急事件检出率95.8%,高于所有医生,且误报率可控。
- 适合需要大规模远程医疗监控的医院或健康机构使用。
远程患者监测(RPM)产生海量数据,但以往试验因数据过载失败。尽管TIM-HF2显示全天候医生监控可降低30%死亡率,但成本过高难以推广。我们开发了Sentinel AI代理,基于模型上下文协议(MCP),利用21种临床工具和多步推理对RPM生命体征进行上下文分诊。评估包括:(1) 自一致性测试(100次读数×5轮);(2) 与规则阈值对比;(3) 6名临床专家(3名医生,3名护士)通过关联矩阵验证。留一法(LOO)分析显示,该代理在紧急事件敏感性(97.5%)和可操作警报敏感性(90.9%)上均优于每位医生(分别为60.0%和69.5%)。虽存在轻微过度分诊(22.5%),但严重误判案例经独立医生审核后,88-94%被证实为合理升级;共识一致率达100%。自一致性极高(kappa=0.850),单次分诊中位成本仅0.34美元。结论:Sentinel以超越个体医生的敏感度实现高效分诊,通过自动化上下文整合,解决过往RPM试验的核心瓶颈,为实现可扩展的高密度监控提供可行路径。
原文摘要 · Abstract (English)
Background: Remote patient monitoring (RPM) generates vast data, yet landmark trials (Tele-HF, BEAT-HF) failed because data volume overwhelmed clinical staff. While TIM-HF2 showed 24/7 physician-led monitoring reduces mortality by 30%, this model remains prohibitively expensive and unscalable. Methods: We developed Sentinel, an autonomous AI agent using Model Context Protocol (MCP) for contextual triage of RPM vitals via 21 clinical tools and multi-step reasoning. Evaluation included: (1) self-consistency (100 readings x 5 runs); (2) comparison against rule-based thresholds; and (3) validation against 6 clinicians (3 physicians, 3 NPs) using a connected matrix design. A leave-one-out (LOO) analysis compared the agent against individual clinicians; severe overtriage cases underwent independent physician adjudication. Results: Against a human majority-vote standard (N=467), the agent achieved 95.8% emergency sensitivity and 88.5% sensitivity for all actionable alerts (85.7% specificity). Four-level exact accuracy was 69.4% (quadratic-weighted kappa=0.778); 95.9% of classifications were within one severity level. In LOO analysis, the agent outperformed every clinician in emergency sensitivity (97.5% vs. 60.0% aggregate) and actionable sensitivity (90.9% vs. 69.5%). While disagreements skewed toward overtriage (22.5%), independent adjudication of severe gaps (>=2 levels) validated agent escalation in 88-94% of cases; consensus resolution validated 100%. The agent showed near-perfect self-consistency (kappa=0.850). Median cost was $0.34/triage. Conclusions: Sentinel triages RPM vitals with sensitivity exceeding individual clinicians. By automating systematic context synthesis, Sentinel addresses the core limitation of prior RPM trials, offering a scalable path toward the intensive monitoring shown to reduce mortality while maintaining a clinically defensible overtriage profile.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。