用大模型生成临床文本,提升皮肤病多模态模型性能。
Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI
- 设计提示词与医学元数据注入策略生成合成临床笔记。
- 在跨域场景下分类准确率显著提升,且实现跨模态检索能力。
- 适合关注医疗多模态数据增强的开发者与研究者。
多模态(MM)学习在生物医学人工智能中日益重要,通过融合互补模态信息以全面反映患者健康状况。然而,异构生物医学多模态数据稀缺限制了鲁棒模型的发展。例如,在皮肤科领域,常见数据集仅包含图像与少量元数据,难以充分发挥多模态整合优势。近期大语言模型(LLMs)可生成图像发现的文本描述,或有望结合图像与文本表征。但这些模型未针对医学领域训练,直接使用存在幻觉风险。本文研究提示词设计与医学元数据引入策略对生成合成临床笔记的影响,并评估其对多模态架构在分类与跨模态检索任务中的性能提升效果。在多个异构皮肤科数据集上的实验表明,合成临床笔记不仅能提升分类性能,尤其在域偏移情况下,还激活了未显式优化的跨模态检索能力。
原文摘要 · Abstract (English)
Multimodal (MM) learning is emerging as a promising paradigm in biomedical artificial intelligence (AI) applications, integrating complementary modality, which highlight different aspects of patient health. The scarcity of large heterogeneous biomedical MM data has restrained the development of robust models for medical AI applications. In the dermatology domain, for instance, skin lesion datasets typically include only images linked to minimal metadata describing the condition, thereby limiting the benefits of MM data integration for reliable and generalizable predictions. Recent advances in Large Language Models (LLMs) enable the synthesis of textual description of image findings, potentially allowing the combination of image and text representations. However, LLMs are not specifically trained for use in the medical domain, and their naive inclusion has raised concerns about the risk of hallucinations in clinically relevant contexts. This work investigates strategies for generating synthetic textual clinical notes, in terms of prompt design and medical metadata inclusion, and evaluates their impact on MM architectures toward enhancing performance in classification and cross-modal retrieval tasks. Experiments across several heterogeneous dermatology datasets demonstrate that synthetic clinical notes not only enhance classification performance, particularly under domain shift, but also unlock cross-modal retrieval capabilities, a downstream task that is not explicitly optimized during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。