用自动优化提示词实现医疗问答多任务统一解法
Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs
- 分阶段独立处理医疗问答,每阶段用AI优化提示词
- 自一致性投票减少错误,各阶段有专属验证机制
- 无需微调模型,在四项任务中表现优异
电子健康记录(EHR)的自动化问答需精准证据检索、忠实答案生成及答案与临床笔记的明确对齐。本文提出Neural1.5方法,参加CL4Health@LREC 2026的ArchEHR-QA 2026共享任务,涵盖四个子任务:问题理解、证据识别、答案生成和证据对齐。方法将任务分解为独立模块化阶段,采用DSPy的MIPROv2优化器自动发现高性能提示词,联合优化各阶段的指令与少样本示例。每个阶段通过多次随机推理的自一致性投票抑制虚假错误,提升可靠性;同时引入阶段特定验证机制(如自省与链式验证)进一步优化输出质量。在参与全部四子任务的队伍中,整体排名第二(平均排名4.00),各子任务分别位列第4、第1、第4、第7。结果表明,系统性分阶段提示优化结合自一致性机制,是复杂临床问答任务中低成本替代模型微调的有效方案。
原文摘要 · Abstract (English)
Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit grounding of answers in clinical notes. In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises four subtasks: question interpretation, evidence identification, answer generation, and evidence alignment. Our approach decouples the task into independent, modular stages and employs DSPy"s MIPROv2 optimizer to automatically discover high-performing prompts, jointly tuning instructions and few-shot demonstrations for each stage. Within every stage, self-consistency voting over multiple stochastic inference runs suppresses spurious errors and improves reliability, while stage-specific verification mechanisms (e.g., self-reflection and chain-of-verification for alignment) further refine output quality. Among all teams that participated in all four subtasks, our method ranks second overall (mean rank 4.00), placing 4th, 1st, 4th, and 7th on Subtasks 1-4, respectively. These results demonstrate that systematic, per-stage prompt optimization combined with self-consistency mechanisms is a cost-effective alternative to model fine-tuning for multifaceted clinical QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。