arXiv:2509.13590eess.IVcs.AI2025-09被引 1

用视觉语言模型自动分析医学影像并生成报告,提升诊断效率。

Intelligent Healthcare Imaging Platform: A VLM-Based Framework for Automated Medical Image Analysis and Clinical Report Generation

  • 基于VLM融合图像与文本,实现跨模态智能诊断
  • 在多模态影像中检测异常,定位误差平均80像素
  • 零样本学习适配临床场景,适合医疗开发者与放射科医生

人工智能在医疗影像领域的快速发展正重塑诊断医学与临床决策。本文提出一种基于视觉语言模型(VLM)的智能多模态医学影像分析框架,集成Google Gemini 2.5 Flash,实现对CT、MRI、X光和超声等多模态影像的自动肿瘤检测与临床报告生成。系统结合视觉特征提取与自然语言处理,支持上下文化图像解读,采用坐标验证机制与概率高斯建模分析异常分布。通过多层可视化技术生成详细医学图示、对比叠加与统计图表,提升临床可信度,定位测量平均偏差为80像素。结果处理采用精准提示工程与文本分析,提取结构化临床信息并保持可解释性。实验表明该系统在多模态异常检测中表现优异。系统配备用户友好的Gradio界面,支持临床工作流集成,具备零样本学习能力,降低对大规模数据集的依赖。该框架显著提升自动化诊断支持与放射科工作效率,但尚需临床验证与多中心评估方可广泛推广。

原文摘要 · Abstract (English)

The rapid advancement of artificial intelligence (AI) in healthcare imaging has revolutionized diagnostic medicine and clinical decision-making processes. This work presents an intelligent multimodal framework for medical image analysis that leverages Vision-Language Models (VLMs) in healthcare diagnostics. The framework integrates Google Gemini 2.5 Flash for automated tumor detection and clinical report generation across multiple imaging modalities including CT, MRI, X-ray, and Ultrasound. The system combines visual feature extraction with natural language processing to enable contextual image interpretation, incorporating coordinate verification mechanisms and probabilistic Gaussian modeling for anomaly distribution. Multi-layered visualization techniques generate detailed medical illustrations, overlay comparisons, and statistical representations to enhance clinical confidence, with location measurement achieving 80 pixels average deviation. Result processing utilizes precise prompt engineering and textual analysis to extract structured clinical information while maintaining interpretability. Experimental evaluations demonstrated high performance in anomaly detection across multiple modalities. The system features a user-friendly Gradio interface for clinical workflow integration and demonstrates zero-shot learning capabilities to reduce dependence on large datasets. This framework represents a significant advancement in automated diagnostic support and radiological workflow efficiency, though clinical validation and multi-center evaluation are necessary prior to widespread adoption.

医学影像视觉语言模型报告生成AI诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。