用医生社交平台数据构建百万级医学图文问答集,训练出更懂临床逻辑的AI模型。
A visual large language foundational model for medical image recognition using clinician-oriented social media
- 从医生社交平台收集去标识化图像与专家评论,构建高质量医学图文数据集。
- 新模型在42个基准上达85.4%宏准确率,比现有模型高3-5%且回答更符合临床思维。
- 适合医学AI研究者、临床辅助系统开发者,推动真实场景下的多模态模型发展。
大型语言模型(LLM)在多个领域展现出强大能力,在医学领域也具有巨大潜力。然而,其在医疗场景中的应用受限于缺乏能体现临床推理和明确图像-文本对齐的视觉问答(VQA)数据集。本文利用去标识化的医学图像及医生在专业社交平台上的专家评论,结合人工验证流程,构建了包含超过一百万对图像-文本问答的长篇医学VQA数据集ThoughtMed-1M,旨在捕捉结构化的临床逻辑与医学图像-文本对齐关系。为验证其有效性,我们开发了基于ThoughtMed-1M训练的奠基型多模态大模型FOLTMed。该模型在42个医学VQA基准测试中达到85.4%的宏准确率,优于当前最优模型3–5%,并在ThoughtMed-1M测试集上生成更具临床一致性的回答,证明了一种可扩展的临床驱动多模态大模型研究范式。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。