用多视角视频提升心脏超声自动解读准确率
Video CLIP Model for Multi-View Echocardiography Interpretation
- 基于五个标准视角的完整视频输入,而非单帧图像
- 在6万+视频-报告对上训练,显著提升检索性能
- 适合心脏病学研究者与医学AI开发者使用
超声心动图通过心脏的超声视频记录心功能信息。近期大规模视觉-语言模型(VLM)的发展推动了超声心动图自动化解读的研究。然而,现有医疗VLM大多依赖单帧(图像)输入,对于仅通过心脏运动才能识别的病症会降低诊断准确性。此外,超声心动图从多个视角采集,不同视角对特定疾病的检测能力各异。利用多视角信息可能提升诊断性能。我们开发了一种视频-语言模型,可处理五个标准视角的完整视频序列,并在60,747对超声心动图视频与报告数据上进行训练。我们评估了视频输入和多视角支持带来的检索性能提升,以及不同预训练模型的贡献。代码与模型权重已在https://github.com/UTcardiology/video-echo-clip公开。
原文摘要 · Abstract (English)
Echocardiography records ultrasound videos of the heart, enabling clinicians to assess cardiac function. Recent advances in large-scale vision-language models (VLMs) have spurred interest in automating echocardiographic interpretation. However, most existing medical VLMs rely on single-frame (image) inputs, which can reduce diagnostic accuracy for conditions identifiable only through cardiac motion. In addition, echocardiographic videos are captured from multiple views, each varying in suitability for detecting specific conditions. Leveraging multiple views may therefore improve diagnostic performance. We developed a video-language model that processes full video sequences from five standard views, trained on 60,747 echocardiographic video-report pairs. We evaluated the gains in retrieval performance from video input and multi-view support, including the contributions of various pretrained models. Code and model weights are available at https://github.com/UTcardiology/video-echo-clip
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。