arXiv:2603.22179cs.AI2026-03

MARCUS用智能体架构实现心电图等多模态心脏影像的自动诊断,准确率超现有模型。

MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management

  • 采用分层智能体架构,融合多模态医学影像与语言模型进行自主推理。
  • 在心电图、超声心动图和核磁共振上分别达到87%-91%、67%-86%、85%-88%准确率。
  • 支持多模态联合分析,且抗幻觉能力强,适合临床辅助决策场景。

心血管疾病仍是全球首要死亡原因,其进展受限于复杂心脏检测的人工解读。现有AI视觉语言模型仅支持单模态输入且不具备交互性。本文提出MARCUS(Multimodal Autonomous Reasoning and Chat for Ultrasound and Signals),一个端到端的多模态智能体系统,可独立或联合处理心电图(ECGs)、超声心动图及心脏磁共振成像(CMR)。MARCUS采用分层智能体架构,包含各模态专用的视觉语言专家模型,每个模型整合领域训练的视觉编码器与多阶段语言模型优化,并由多模态协调器统一调度。模型基于1350万张图像(25万份心电图、130万张超声图像、1200万张CMR图像)和160万条专家标注问题数据集训练。在斯坦福(Stanford)与旧金山大学(UCSF)内部及外部测试集上,心电图准确率达87-91%,超声心动图达67-86%,CMR达85-88%,优于前沿模型(如GPT-5 Thinking、Gemini 2.5 Pro Deep Think)34-45%(P<0.001)。多模态联合分析准确率达70%,接近前沿模型的三倍(22-28%),自由文本质量得分提升1.7-3.0倍。该架构还具备对抗‘幻觉推理’的能力,防止模型从非预期文本信号或虚构视觉内容中推导错误结论。研究证明,结合领域视觉编码器与智能体协调机制,可实现高效可靠的多模态心脏诊疗。模型、代码与基准已开源。

原文摘要 · Abstract (English)

Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Current AI vision-language models are limited to single-modality inputs and are non-interactive. We present MARCUS (Multimodal Autonomous Reasoning and Chat for Ultrasound and Signals), an agentic vision-language system for end-to-end interpretation of electrocardiograms (ECGs), echocardiograms, and cardiac magnetic resonance imaging (CMR) independently and as multimodal input. MARCUS employs a hierarchical agentic architecture comprising modality-specific vision-language expert models, each integrating domain-trained visual encoders with multi-stage language model optimization, coordinated by a multimodal orchestrator. Trained on 13.5 million images (0.25M ECGs, 1.3M echocardiogram images, 12M CMR images) and our novel expert-curated dataset spanning 1.6 million questions, MARCUS achieves state-of-the-art performance surpassing frontier models (GPT-5 Thinking, Gemini 2.5 Pro Deep Think). Across internal (Stanford) and external (UCSF) test cohorts, MARCUS achieves accuracies of 87-91% for ECG, 67-86% for echocardiography, and 85-88% for CMR, outperforming frontier models by 34-45% (P<0.001). On multimodal cases, MARCUS achieved 70% accuracy, nearly triple that of frontier models (22-28%), with 1.7-3.0x higher free-text quality scores. Our agentic architecture also confers resistance to mirage reasoning, whereby vision-language models derive reasoning from unintended textual signals or hallucinated visual content. MARCUS demonstrates that domain-specific visual encoders with an agentic orchestrator enable multimodal cardiac interpretation. We release our models, code, and benchmark open-source.

多模态医疗AI智能体心脏病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。