用可解释的自然语言提示集提升医学视觉语言模型诊断可信度
BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models
- 用大模型自动生成多样且可读的提示对,替代传统不可解释的向量
- 在少样本场景下超越现有方法,提示与临床显著特征高度语义匹配
- 适合需要可解释性医疗AI的临床研究者和开发者
医学视觉语言模型的临床应用受限于提示优化技术产生的不可解释隐向量或单一文本提示。这种不透明性及未能捕捉临床诊断中多维度观察整合的本质,降低了其在高风险场景中的可信度。为此,我们提出BiomedXPro,一种进化式框架,利用大语言模型作为生物医学知识提取器和自适应优化器,自动生成一组多样、可解释、自然语言的提示对用于疾病诊断。在多个生物医学基准测试中,BiomedXPro持续优于最先进的提示调优方法,尤其在数据稀缺的少样本设置下表现突出。进一步分析显示,发现的提示与统计显著的临床特征具有强语义对齐,使模型性能建立在可验证的概念基础上。通过生成多样且可解释的提示集合,BiomedXPro为模型预测提供了可验证依据,是构建更可信、更符合临床需求的AI系统的关键一步。
原文摘要 · Abstract (English)
The clinical adoption of biomedical vision-language models is hindered by prompt optimization techniques that produce either uninterpretable latent vectors or single textual prompts. This lack of transparency and failure to capture the multi-faceted nature of clinical diagnosis, which relies on integrating diverse observations, limits their trustworthiness in high-stakes settings. To address this, we introduce BiomedXPro, an evolutionary framework that leverages a large language model as both a biomedical knowledge extractor and an adaptive optimizer to automatically generate a diverse ensemble of interpretable, natural-language prompt pairs for disease diagnosis. Experiments on multiple biomedical benchmarks show that BiomedXPro consistently outperforms state-of-the-art prompt-tuning methods, particularly in data-scarce few-shot settings. Furthermore, our analysis demonstrates a strong semantic alignment between the discovered prompts and statistically significant clinical features, grounding the model's performance in verifiable concepts. By producing a diverse ensemble of interpretable prompts, BiomedXPro provides a verifiable basis for model predictions, representing a critical step toward the development of more trustworthy and clinically-aligned AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。