多智能体系统自动评估心律失常,给出可审计的临床决策建议。
Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

- 每个生理信号由独立智能体分析,结合患者数据与指南推理。
- 在诊断、意义、紧急程度和管理上均优于专家和大模型基准。
- 决策可追溯至引用指南,适合用于连续心电监测场景。
同一心房颤动事件在健康成人中可能是轻微发现,而在高血压老年患者中则需抗凝治疗:相同信号,不同判断。命名心律只是起点;决定患者预后的,是后续的综合判断——心律在整个记录中的表现、对患者的含义及应对措施。现有将大语言模型与心电图结合的研究止步于单次记录解读,未能实现患者层面整合;而基于此构建的智能体系统或依赖设备已检出的心律异常,或服务于其他诊断任务,未触及最终临床决策。本文提出 Cardiologent,一个覆盖从检测到决策全过程的多智能体系统。每个信号(单导联心电图与可穿戴设备获取的光电容积脉搏波)由独立智能体分析,依据测量特征而非单纯标签;各片段结果整合为患者心律谱,并结合患者自身数据,参照针对该病例检索到的临床指南进行推理,由审核智能体验证每项结论是否符合其引用指南。评估聚焦于临床决策本身,涵盖整合诊断、临床意义、紧急程度与管理策略。Cardiologent 在所有维度得分最高,在多位心脏病专家与大规模语言模型裁判下均排名第一,两者一致性(ICC 0.74, 0.66)接近专家间一致(0.67)。因每项结论均可追溯至所引指南且经专家验证,系统输出结果具备可审计性,推动其在持续监测中的实际应用。
原文摘要 · Abstract (English)
The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language models with the ECG stops short of this, reading one recording without assembling a patient-level finding; and agentic systems built around it either receive the arrhythmia a device has already detected or target a different diagnostic task, stopping before the decision this task requires. We formulate patient-level arrhythmia decision support as a task and present Cardiologent, a multi-agent system that spans it from detection to decision. An agent for each signal -- a single ECG lead and the photoplethysmogram a wearable acquires -- grounds its window reading in measured features rather than a bare label; the readings are assembled into the patient's rhythm profile and, with the patient's own data, reasoned against clinical guidelines retrieved for the case, with a critic checking each conclusion against the guideline it cites. We evaluate the clinical decision rather than the report, across integrated diagnosis, clinical significance, and urgency and management. Cardiologent scores highest on every axis, first on every patient-level task under both cardiologists and an at-scale LLM judge -- whose agreement with the cardiologists (ICC 0.74, 0.66) matches theirs with each other (0.67). Because each conclusion traces to a cited guideline and is validated against expert cardiologists, it yields decisions a clinician can audit rather than act on blindly -- a step toward use in continuous monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。