arXiv:2502.13447cs.CVcs.CL2025-02中稿 · ICASSP'25

用医学知识增强跨模态学习,提升胸部X光分类准确率

Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning

  • 基于集合论生成可控粒度的医学描述文本注入模型
  • 细粒度知识注入使准确率提升至72.5%,远超人工描述的49.9%
  • 适合医疗影像分析与多模态模型优化的研究者参考

人工智能在医学影像中的应用潜力巨大,但预训练知识与跨模态学习性能之间的关系尚不明确。本研究探讨在跨模态分类中显式注入医学知识对胸部X光(CXR)图像分类的影响。提出一种基于集合论的知识注入框架,可生成具有可控知识粒度的CXR图像描述。利用该框架,我们在包含不同医学信息量的描述上微调CLIP模型,并在CheXpert数据集上通过零样本分类评估性能。结果表明,注入细粒度医学知识显著提升分类准确率,达到72.5%,而使用人工生成描述时仅为49.9%。此外,我们发现知识密度更高及采用领域专用大语言模型(LLM)生成描述,有助于进一步提升性能。本研究验证了知识注入在提升自动化胸部X光分类中的有效性,为更精准可靠的诊断工具提供新路径。

原文摘要 · Abstract (English)

The integration of artificial intelligence in medical imaging has shown tremendous potential, yet the relationship between pre-trained knowledge and performance in cross-modality learning remains unclear. This study investigates how explicitly injecting medical knowledge into the learning process affects the performance of cross-modality classification, focusing on Chest X-ray (CXR) images. We introduce a novel Set Theory-based knowledge injection framework that generates captions for CXR images with controllable knowledge granularity. Using this framework, we fine-tune CLIP model on captions with varying levels of medical information. We evaluate the model's performance through zero-shot classification on the CheXpert dataset, a benchmark for CXR classification. Our results demonstrate that injecting fine-grained medical knowledge substantially improves classification accuracy, achieving 72.5\% compared to 49.9\% when using human-generated captions. This highlights the crucial role of domain-specific knowledge in medical cross-modality learning. Furthermore, we explore the influence of knowledge density and the use of domain-specific Large Language Models (LLMs) for caption generation, finding that denser knowledge and specialized LLMs contribute to enhanced performance. This research advances medical image analysis by demonstrating the effectiveness of knowledge injection for improving automated CXR classification, paving the way for more accurate and reliable diagnostic tools.

医学影像知识注入跨模态学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。