多视角推理模型提升心脏超声诊断准确率与报告可信度
EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography

- 融合时空编码与解剖定位,实现跨视角医学图像推理
- 在私有数据集上分类准确率提升17.1%,临床报告可信度达0.800
- 适合需要可解释性诊断辅助的临床场景
超声心动图是最广泛使用的无创心脏影像技术,提供心血管诊断的关键信息。解读超声心动图需整合多个心脏视角的互补证据以识别异常并生成结构化临床报告。尽管近期研究聚焦于提升分类性能,多数模型缺乏显式的诊断推理和空间定位的解剖证据,限制了临床信任。我们提出EchoSonar-R,一种多视角推理增强的视觉语言模型,可联合完成超声心动图的多标签疾病分类与报告生成。EchoSonar-R结合时空视频编码器与结构感知的心脏检测器,提供空间定位的解剖线索,提升跨视角推理的可解释性与临床可信度。模型采用两阶段训练:首先在带推理标注的目标上进行监督微调(SFT),随后通过任务特异性奖励的组相对策略优化(GRPO)进行强化学习联合对齐分类与报告生成。在私有多视角数据集及两个公开基准上,EchoSonar-R在私有集上宏平衡准确率提升17.1%,在MIMICEchoQA上提升6.1%,绿色临床忠实度评分为0.800,并生成基于多视角视觉证据的可解释推理路径。
原文摘要 · Abstract (English)
Echocardiography is the most widely used non-invasive cardiac imaging modality, providing essential information for cardiovascular diagnosis. Interpreting an echocardiogram requires synthesizing complementary evidence across multiple heart views to identify abnormalities and produce structured clinical reports. While recent efforts focus on improving classification performance, most models lack explicit diagnostic reasoning and spatially grounded anatomical evidence, limiting clinician trust. We present EchoSonar-R, a multi-view reasoning-enabled vision-language model that jointly performs multi-label disease classification and report generation from echocardiography studies. EchoSonar-R combines a spatiotemporal video encoder with a structure-aware cardiac detector that provides spatially grounded anatomical cues to improve interpretability and clinician trust during cross-view reasoning. EchoSonar-R is trained in two stages: supervised fine-tuning (SFT) on reasoning-annotated targets, followed by Group Relative Policy Optimization (GRPO) with task-specific rewards that jointly align classification and report generation within a unified reinforcement-learning framework. Across a private multi-view dataset and two public benchmarks, EchoSonar-R improves macro balanced accuracy by 17.1% on the private set and 6.1% on MIMICEchoQA over the strongest baseline, achieves a GREEN clinical faithfulness score of 0.800, and produces interpretable reasoning traces grounded in multi-view visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。