构建医院内网的隔离式大模型系统,保障医疗数据安全并支持临床实用。
Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation

- 采用容器化隔离架构,强制网络分段与出站过滤,防止数据外泄。
- 10名医生试用一周,文本修正类任务获高临床评价,开放生成易出错。
- 开源部署工具包,适合需合规部署LLM的医疗机构参考。
目的:设计、实现、评估并报告符合最小权限原则的自托管大语言模型(LLM)基础设施在放射科的应用,强调技术可行性、网络隔离和临床实用性。方法:采用隔离优先的容器化推理栈,实施严格的网络分割、主机强制出站过滤及主动隔离监控,防止未经授权的外部连接。配套部署包提供自动化隔离与加固测试。系统通过vLLM服务开源权重的DeepSeek-R1模型。在为期一周的试点阶段,22名住院医师和放射科医生可自由使用10个预设提示模板,在日常工作中按需使用。之后,他们对临床实用性与系统稳定性进行0-10分评分,并报告模型输出中的关键错误。结果:该机构治理路径获得临床管理、合规、数据保护与信息安全负责人批准,可处理未匿名化的受保护健康信息(PHI)。试点期间系统稳定且用户友好。基于源文本的任务(如报告修改或简化、指南推荐)获最高实用性评分;而基于影像发现的开放式结论生成则出现最高频的关键错误,如具有临床意义的幻觉或遗漏。结论:所提出的隔离优先本地部署架构成功跨越监管障碍,显示出在文本锚定任务中的良好临床价值,现已成为德国一所超万名员工大学医院正式提供开源权重LLM服务的基础。部署包已公开(https://github.com/ukbonn/ukb-gpt)。
原文摘要 · Abstract (English)
Purpose: To design, implement, evaluate, and report on the regulatory requirements of a self-hosted LLM infrastructure for radiology adhering to the principle of least privilege, emphasizing technical feasibility, network isolation, and clinical utility. Materials and Methods: The isolation-first, containerized LLM inference stack relies on strict network segmentation, host-enforced egress filtering, and active isolation monitoring preventing unauthorized external connectivity. An accompanying deployment package provides automated isolation and hardening tests. The system served the open-weights DeepSeek-R1 model via vLLM. In a one-week pilot phase, 22 residents and radiologists were free to use 10 predefined prompt-templates whenever they considered them useful in daily work. Afterward, they rated clinical utility and system stability on an 0-10 Likert scale and reported observed critical errors in model output. Results: The applied institutional governance pathway achieved approval from clinic management, compliance, data protection and information security officers for processing unanonymized PHI. The system was rated stable and user friendly during the pilot. Source text-anchored tasks, such as report corrections or simplifications, and radiology guideline recommendations received the highest utility ratings, whereas open-ended conclusion generation based on findings resulted in the highest frequency of critical errors, such as clinically relevant hallucinations or omissions. Conclusion: The proposed isolation-first on-premise architecture enabled overcoming regulatory borders, showed promising clinical utility in text-anchored tasks and is the current base to serve open-weights LLMs as an official service of a German University Hospital with over 10,000 employees. The deployment package were made publicly available (https://github.com/ukbonn/ukb-gpt).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。