arXiv:2601.12174eess.IV2026-01被引 3

多任务视觉语言模型助诊右上腹超声,提升诊断一致性与报告效率。

A multitask framework for automated interpretation of multi-frame right upper quadrant ultrasound in clinical decision support

  • 基于多帧图像与临床文本的联合理解,实现疾病分类、报告生成与手术建议
  • 在三所医院数据集上均达高准确率,报告质量与专家水平无异
  • 适合临床辅助诊断场景,尤其适用于超声资源紧张的医疗环境

超声是急诊和肝胆影像的核心手段,但其解读高度依赖操作者且时间敏感。本文提出一种多任务视觉语言代理(VLM),用于支持全诊疗流程的右上腹(RUQ)超声综合解读。该系统基于Qwen2.5-VL-7B架构,在大规模多中心数据集上训练,主队列来自约翰霍普金斯医学院(9,189例,594,099张图像),并在斯坦福大学(108例,3,240张图像)和中国某大型医学中心(257例,3,178张图像)进行外部验证。模型整合帧级视觉理解与报告引导的语言推理,完成三项任务:(i) 18种肝胆及胆囊疾病的分类,(ii) 生成临床一致的诊断报告,(iii) 基于超声发现与临床数据提供手术决策支持。模型在各项任务中均表现高诊断准确率,生成报告在盲评中与专家撰写版本无差异,内容型评估显示其事实准确性和信息密度更优。该代理能以高精度识别需行胆囊切除术患者,支持实时决策。结果表明,通用视觉语言模型有望提升超声诊断的一致性、报告效率与手术分诊能力。

原文摘要 · Abstract (English)

Ultrasound is a cornerstone of emergency and hepatobiliary imaging, yet its interpretation remains highly operator-dependent and time-sensitive. Here, we present a multitask vision-language agent (VLM) developed to assist with comprehensive right upper quadrant (RUQ) ultrasound interpretation across the full diagnostic workflow. The system was trained on a large, multi-center dataset comprising a primary cohort from Johns Hopkins Medical Institutions (9,189 cases, 594,099 images) and externally validated on cohorts from Stanford University (108 cases, 3,240 images) and a major Chinese medical center (257 cases, 3,178 images). Built on the Qwen2.5-VL-7B architecture, the agent integrates frame-level visual understanding with report-grounded language reasoning to perform three tasks: (i) classification of 18 hepatobiliary and gallbladder conditions, (ii) generation of clinically coherent diagnostic reports, and (iii) surgical decision support based on ultrasound findings and clinical data. The model achieved high diagnostic accuracy across all tasks, generated reports that were indistinguishable from expert-written versions in blinded evaluations, and demonstrated superior factual accuracy and information density on content-based metrics. The agent further identified patients requiring cholecystectomy with high precision, supporting real-time decision-making. These results highlight the potential of generalist vision-language models to improve diagnostic consistency, reporting efficiency, and surgical triage in real-world ultrasound practice.

超声辅助视觉语言模型临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。