用影像+病历联合诊断肺结核,准确率超96%。
Vision-Language Models for Acute Tuberculosis Diagnosis: A Multimodal Approach Combining Imaging and Clinical Data
- 融合胸部X光与临床文本,用视觉语言模型生成诊断报告
- 对实变、空洞等关键病灶检测精度达97%,召回率达96%
- 适合资源有限地区快速筛查,减轻医生负担
本研究提出一种基于SIGLIP和Gemma-3b架构的视觉语言模型(VLM),用于自动化急性肺结核(TB)筛查。通过整合胸部X光图像与临床病历,模型旨在提升诊断准确性和效率,尤其适用于资源匮乏地区。该模型采用SIGLIP进行图像编码,Gemma-3b进行文本解码,有效捕捉急性肺结核特异性病灶与临床信息。结果显示,对实变、空洞、结节等关键病理特征的检测精度为97%,召回率为96%。模型具备良好的空间定位能力,能可靠区分肺结核阳性病例。多模态能力降低了对放射科医生的依赖,为急性肺结核筛查提供了可扩展的解决方案。未来工作将聚焦于提升对细微病灶的识别能力,并缓解数据集偏差,以增强模型在多样全球医疗环境中的泛化性与适用性。
原文摘要 · Abstract (English)
Background: This study introduces a Vision-Language Model (VLM) leveraging SIGLIP and Gemma-3b architectures for automated acute tuberculosis (TB) screening. By integrating chest X-ray images and clinical notes, the model aims to enhance diagnostic accuracy and efficiency, particularly in resource-limited settings. Methods: The VLM combines visual data from chest X-rays with clinical context to generate detailed, context-aware diagnostic reports. The architecture employs SIGLIP for visual encoding and Gemma-3b for decoding, ensuring effective representation of acute TB-specific pathologies and clinical insights. Results: Key acute TB pathologies, including consolidation, cavities, and nodules, were detected with high precision (97percent) and recall (96percent). The model demonstrated strong spatial localization capabilities and robustness in distinguishing TB-positive cases, making it a reliable tool for acute TB diagnosis. Conclusion: The multimodal capability of the VLM reduces reliance on radiologists, providing a scalable solution for acute TB screening. Future work will focus on improving the detection of subtle pathologies and addressing dataset biases to enhance its generalizability and application in diverse global healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。