arXiv:2511.11898cs.CVcs.AI2025-11

自动化优化提示词,让医疗视觉语言模型性能提升最高达34倍。

Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks

  • 用结构化方法自动优化提示词,不依赖人工设计。
  • 中位性能提升53%,部分任务提升超3000%。
  • 适合希望快速部署医疗AI但缺乏数据的医疗机构。

视觉-语言基础模型在多种影像任务中展现出潜力,但在医疗基准测试中表现常不理想。以往改进方法包括需要大量领域数据和计算资源的微调,或难以泛化的人工提示工程,二者均限制了医疗机构的部署。为此,本文采用声明式自优化Python(DSPy)框架,对五种跨放射科、胃肠科和皮肤科的医学影像任务进行系统性评估,测试四种提示优化技术在十种开源视觉-语言模型上的表现。优化后管道相比零样本提示基线实现中位53%的相对提升,部分任务提升高达300%至3,400%,尤其在原始性能较低的任务上效果显著。结果表明,自动化提示优化能显著提升医疗AI系统的视觉理解能力,减少对人工提示设计的依赖,使临床人员更专注于患者诊疗。该方法具备可扩展性且保护数据隐私,适用于开源模型。我们已公开评估流程,支持可复现研究,详见https://github.com/DaneshjouLab/prompt-triage-lab。

原文摘要 · Abstract (English)

Vision-language foundation models (VLMs) show promise for diverse imaging tasks but often underperform on medical benchmarks. Prior efforts to improve performance include model finetuning, which requires large domain-specific datasets and significant compute, or manual prompt engineering, which is hard to generalize and often inaccessible to medical institutions seeking to deploy these tools. These challenges motivate interest in approaches that draw on a model's embedded knowledge while abstracting away dependence on human-designed prompts to enable scalable, weight-agnostic performance improvements. To explore this, we adapt the Declarative Self-improving Python (DSPy) framework for structured automated prompt optimization in medical vision-language systems through a comprehensive, formal evaluation. We implement prompting pipelines for five medical imaging tasks across radiology, gastroenterology, and dermatology, evaluating 10 open-source VLMs with four prompt optimization techniques. Optimized pipelines achieved a median relative improvement of 53% over zero-shot prompting baselines, with the largest gains ranging from 300% to 3,400% on tasks where zero-shot performance is low. These results highlight the substantial potential of applying automated prompt optimization to medical AI systems, demonstrating significant gains for vision-based applications requiring accurate clinical image interpretation. By reducing dependence on prompt design to elicit intended outputs, these techniques allow clinicians to focus on patient care and clinical decision-making. Furthermore, our experiments offer scalability and preserve data privacy, demonstrating performance improvement on open-source VLMs. We publicly release our evaluation pipelines to support reproducible research on specialized medical tasks, available at https://github.com/DaneshjouLab/prompt-triage-lab.

视觉语言模型医疗AI提示优化自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。