优化临床大模型提示词,让回答更稳定可靠。
Stability-Aware Prompt Optimization for Clinical Data Abstraction
- 设计双目标优化循环,同时提升准确率与提示稳定性。
- 在多个模型上测试,提示词改写后翻转率显著下降。
- 适合医疗领域大模型验证与部署的开发者参考。
用于临床数据抽象的大语言模型对提示词措辞敏感,但现有研究多将提示词视为固定,孤立分析不确定性。我们提出应联合考虑提示词与模型表现。在两个临床任务(MedAlign适用性/正确性、MS亚型抽象)及多个开源与专有模型中,通过翻转率衡量提示敏感性,并关联校准性与选择性预测。结果表明,高准确率不意味着提示稳定,模型可能看似校准良好却对改写极为脆弱。我们提出一种双目标提示优化循环,同时优化准确率与稳定性,实验证明显式引入稳定性项可降低各任务与模型的翻转率,有时仅以轻微准确率损失为代价。结果表明,临床大模型系统验证时应将提示敏感性作为明确目标。
原文摘要 · Abstract (English)
Large language models used for clinical abstraction are sensitive to prompt wording, yet most work treats prompts as fixed and studies uncertainty in isolation. We argue these should be treated jointly. Across two clinical tasks (MedAlign applicability/correctness and MS subtype abstraction) and multiple open and proprietary models, we measure prompt sensitivity via flip rates and relate it to calibration and selective prediction. We find that higher accuracy does not guarantee prompt stability, and that models can appear well-calibrated yet remain fragile to paraphrases. We propose a dual-objective prompt optimization loop that jointly targets accuracy and stability, showing that explicitly including a stability term reduces flip rates across tasks and models, sometimes at modest accuracy cost. Our results suggest prompt sensitivity should be an explicit objective when validating clinical LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。