用强化学习优化医疗提示词,让大模型更准确安全。
EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
- 设计医学术语注意力机制和多维度评估体系
- 事实错误减少24.7%,领域特异性提升19.6%
- 适合临床提示工程与医疗AI安全研究者
提示词工程显著影响大语言模型在医疗应用中的可靠性与临床实用性。现有优化方法难以兼顾领域专业知识与安全性要求。本文提出EMPOWER,一种新型进化框架,通过专用表示学习、多维评估与结构保持算法提升医疗提示质量。方法包括:(1) 医学术语注意力机制,(2) 评估清晰度、特异性、临床相关性与事实准确性在内的综合评估架构,(3) 保留临床推理完整性的组件级进化算法,(4) 确保符合医学知识的语义验证模块。在诊断、治疗与教育任务中评估显示:事实错误内容减少24.7%,领域特异性提升19.6%,盲评中医生偏好度提高15.3%。该框架解决了临床提示构建的关键挑战,推动大模型更负责任地融入医疗场景。
原文摘要 · Abstract (English)
Prompt engineering significantly influences the reliability and clinical utility of Large Language Models (LLMs) in medical applications. Current optimization approaches inadequately address domain-specific medical knowledge and safety requirements. This paper introduces EMPOWER, a novel evolutionary framework that enhances medical prompt quality through specialized representation learning, multi-dimensional evaluation, and structure-preserving algorithms. Our methodology incorporates: (1) a medical terminology attention mechanism, (2) a comprehensive assessment architecture evaluating clarity, specificity, clinical relevance, and factual accuracy, (3) a component-level evolutionary algorithm preserving clinical reasoning integrity, and (4) a semantic verification module ensuring adherence to medical knowledge. Evaluation across diagnostic, therapeutic, and educational tasks demonstrates significant improvements: 24.7% reduction in factually incorrect content, 19.6% enhancement in domain specificity, and 15.3% higher clinician preference in blinded evaluations. The framework addresses critical challenges in developing clinically appropriate prompts, facilitating more responsible integration of LLMs into healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。