arXiv:2508.20345cs.CVcs.HC2025-08

轻量安全工具箱,让医生轻松用医学视觉语言模型。

MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models

  • 图形化界面支持无编程选型与部署医疗多模态模型
  • 单块A6000 GPU即可本地运行,保障数据隐私
  • 适合临床研究者和工程师快速落地医疗AI应用

近期医学视觉语言模型(VLMs)的发展为自动化报告生成、医生辅助系统和不确定性量化等临床应用带来了巨大机遇。然而,这些模型也带来严重安全风险,如受保护健康信息(PHI)泄露、数据外泄及网络威胁,尤其在医院环境中尤为关键。即使用于科研或非临床场景,医疗机构也需谨慎并采取防护措施。为此,我们提出MedFoundationHub,一个图形化用户界面工具包,具备三大功能:(1) 使医生无需编程即可手动选择并使用不同模型;(2) 支持工程师以即插即用方式高效部署医疗VLM,无缝集成Hugging Face开源模型;(3) 通过容器化(Docker)实现跨操作系统、隐私保护的推理部署。该工具仅需配备单张NVIDIA A6000 GPU的离线本地工作站,确保安全且适用于典型学术实验室资源。为评估能力,我们邀请注册病理学家部署并评估五种前沿VLM(Google-MedGemma3-4B、Qwen2-VL-7B-Instruct、Qwen2.5-VL-7B-Instruct、LLaVA-1.5-7B/13B),覆盖结肠和肾脏病例,共完成1015次医患-模型评分事件。评估揭示常见问题包括答非所问、推理模糊和病理术语不一致。

原文摘要 · Abstract (English)

Recent advances in medical vision-language models (VLMs) open up remarkable opportunities for clinical applications such as automated report generation, copilots for physicians, and uncertainty quantification. However, despite their promise, medical VLMs introduce serious security concerns, most notably risks of Protected Health Information (PHI) exposure, data leakage, and vulnerability to cyberthreats - which are especially critical in hospital environments. Even when adopted for research or non-clinical purposes, healthcare organizations must exercise caution and implement safeguards. To address these challenges, we present MedFoundationHub, a graphical user interface (GUI) toolkit that: (1) enables physicians to manually select and use different models without programming expertise, (2) supports engineers in efficiently deploying medical VLMs in a plug-and-play fashion, with seamless integration of Hugging Face open-source models, and (3) ensures privacy-preserving inference through Docker-orchestrated, operating system agnostic deployment. MedFoundationHub requires only an offline local workstation equipped with a single NVIDIA A6000 GPU, making it both secure and accessible within the typical resources of academic research labs. To evaluate current capabilities, we engaged board-certified pathologists to deploy and assess five state-of-the-art VLMs (Google-MedGemma3-4B, Qwen2-VL-7B-Instruct, Qwen2.5-VL-7B-Instruct, and LLaVA-1.5-7B/13B). Expert evaluation covered colon cases and renal cases, yielding 1015 clinician-model scoring events. These assessments revealed recurring limitations, including off-target answers, vague reasoning, and inconsistent pathology terminology.

医学多模态隐私安全工具链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。