用视觉语言模型融合X光片和病历,提升慢性结核诊断准确率。
Advancing Chronic Tuberculosis Diagnostics Using Vision-Language Models: A Multi modal Framework for Precision Analysis
- 用ViT+Gemma-3b模型融合影像与临床数据,实现跨模态分析。
- 检测关键病变的精确率和召回率均达94%,AUC超0.93。
- 适合资源有限地区使用,可生成带上下文的诊断报告。
本研究提出一种基于视觉语言模型(VLM)的自动化慢性结核(TB)筛查框架,采用SIGLIP编码器与Gemma-3b解码器,融合胸部X光图像与临床信息。模型通过视觉变压器(ViT)处理影像,以变压器文本编码器解析病史等临床背景,利用跨模态注意力对齐图像特征与文本信息,并由Gemma-3b生成完整诊断报告。模型在500万对医学图文数据上预训练,再在10万张慢性结核特异性胸片上微调。结果表明,该模型对纤维化、钙化肉芽肿及支气管扩张等慢性结核病灶的检测精确率与召回率均为94%,曲线下面积(AUC)超过0.93,交并比(IoU)高于0.91,验证了其在病灶识别与定位上的有效性。该模型为自动化慢性结核诊断提供了一种鲁棒且可扩展的解决方案,能整合影像与临床数据,输出具备上下文意识的可操作洞察。未来工作将聚焦于细微病灶识别与数据集偏差问题,以提升模型在不同人群与医疗环境中的泛化能力。
原文摘要 · Abstract (English)
Background: This study proposes a Vision-Language Model (VLM) leveraging the SIGLIP encoder and Gemma-3b transformer decoder to enhance automated chronic tuberculosis (TB) screening. By integrating chest X-ray images with clinical data, the model addresses the challenges of manual interpretation, improving diagnostic consistency and accessibility, particularly in resource-constrained settings. Methods: The VLM architecture combines a Vision Transformer (ViT) for visual encoding and a transformer-based text encoder to process clinical context, such as patient histories and treatment records. Cross-modal attention mechanisms align radiographic features with textual information, while the Gemma-3b decoder generates comprehensive diagnostic reports. The model was pre-trained on 5 million paired medical images and texts and fine-tuned using 100,000 chronic TB-specific chest X-rays. Results: The model demonstrated high precision (94 percent) and recall (94 percent) for detecting key chronic TB pathologies, including fibrosis, calcified granulomas, and bronchiectasis. Area Under the Curve (AUC) scores exceeded 0.93, and Intersection over Union (IoU) values were above 0.91, validating its effectiveness in detecting and localizing TB-related abnormalities. Conclusion: The VLM offers a robust and scalable solution for automated chronic TB diagnosis, integrating radiographic and clinical data to deliver actionable and context-aware insights. Future work will address subtle pathologies and dataset biases to enhance the model's generalizability, ensuring equitable performance across diverse populations and healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。