arXiv:2510.21228cs.CLcs.HC2025-10被引 2

用知识树+智能体模拟急救调度,提升训练与决策支持效果

DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services

  • 基于临床分类体系构建智能体系统,确保对话符合医学逻辑
  • 94%情况下能正确联系其他相关人员,91%案例提供有效建议
  • 适合急救培训、流程评估及未来实时辅助系统的研发

急救调度是高风险过程,常受来电者情绪、信息模糊和认知负荷影响。本文构建了一个包含32种主诉、6类来电人身份的临床分类体系(基于MIMIC-III数据集)和六阶段呼叫流程,开发了一个基于AutoGen的多智能体系统,包含来电者与调度员智能体。系统通过事实共享库确保交互的临床合理性,防止误导信息。采用混合评估方法:四位医生对100个模拟案例进行评分,评估“指导有效性”与“调度有效性”,并辅以自动语言分析(情感、可读性、礼貌度)。结果显示,人工评价具较高一致性(Gwe's AC1 > 0.70),系统表现优异:调度有效性达94%(正确联系潜在关联方),指导有效性为91%(提供有效建议)。算法指标验证结果:73.7%中性情感,90.4%中性情绪,可读性得分80.9(Flesch),60.0%表达礼貌,无不礼貌内容。结论表明,该知识树驱动的多智能体系统可高保真模拟多样化急救场景,适用于调度员培训、流程评估,并为实时辅助决策提供基础。

原文摘要 · Abstract (English)

Objective: Emergency medical dispatch (EMD) is a high-stakes process challenged by caller distress, ambiguity, and cognitive load. Large Language Models (LLMs) and Multi-Agent Systems (MAS) offer opportunities to augment dispatchers. This study aimed to develop and evaluate a taxonomy-grounded, LLM-powered multi-agent system for simulating realistic EMD scenarios. Methods: We constructed a clinical taxonomy (32 chief complaints, 6 caller identities from MIMIC-III) and a six-phase call protocol. Using this framework, we developed an AutoGen-based MAS with Caller and Dispatcher Agents. The system grounds interactions in a fact commons to ensure clinical plausibility and mitigate misinformation. We used a hybrid evaluation framework: four physicians assessed 100 simulated cases for "Guidance Efficacy" and "Dispatch Effectiveness," supplemented by automated linguistic analysis (sentiment, readability, politeness). Results: Human evaluation, with substantial inter-rater agreement (Gwe's AC1 > 0.70), confirmed the system's high performance. It demonstrated excellent Dispatch Effectiveness (e.g., 94 % contacting the correct potential other agents) and Guidance Efficacy (advice provided in 91 % of cases), both rated highly by physicians. Algorithmic metrics corroborated these findings, indicating a predominantly neutral affective profile (73.7 % neutral sentiment; 90.4 % neutral emotion), high readability (Flesch 80.9), and a consistently polite style (60.0 % polite; 0 % impolite). Conclusion: Our taxonomy-grounded MAS simulates diverse, clinically plausible dispatch scenarios with high fidelity. Findings support its use for dispatcher training, protocol evaluation, and as a foundation for real-time decision support. This work outlines a pathway for safely integrating advanced AI agents into emergency response workflows.

急救调度多智能体LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。