arXiv:2604.07116cs.CL2026-04

Yale-DM-Lab用多模型投票提升电子病历问答准确率

Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering

论文配图:Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering
图 1 · 摘自论文原文
  • 采用双模型管道与多模型集成,结合少样本提示和投票机制
  • 在开发集上证据对齐最高达88.81微F1,整体表现优于单模型
  • 适合医疗AI研究者,尤其关注病历问答中的推理与对齐问题

我们介绍耶鲁数据医学实验室参与的ArchEHR-QA 2026共享任务系统。该任务研究患者撰写的住院记录问题,包含四个子任务:临床医生解读的问题重述、证据句识别、答案生成及证据-答案对齐。ST1使用Claude Sonnet 4与GPT-4o组成的双模型流水线重述患者问题。ST2-ST4依赖于部署在Azure上的模型集成(o3、GPT-5.2、GPT-5.1、DeepSeek-R1),结合少样本提示与投票策略。实验显示三个主要结论:第一,模型多样性与集成投票显著优于单模型基线;第二,将完整临床答案段落作为额外提示上下文用于对齐;第三,在开发集上,对齐准确率主要受限于推理能力。最佳得分分别为:ST4微F1 88.81,ST2宏F1 65.72,ST3得分为34.01,ST1为33.05。

原文摘要 · Abstract (English)

We describe the Yale-DM-Lab system for the ArchEHR-QA 2026 shared task. The task studies patient-authored questions about hospitalization records and contains four subtasks (ST): clinician-interpreted question reformulation, evidence sentence identification, answer generation, and evidence-answer alignment. ST1 uses a dual-model pipeline with Claude Sonnet 4 and GPT-4o to reformulate patient questions into clinician-interpreted questions. ST2-ST4 rely on Azure-hosted model ensembles (o3, GPT-5.2, GPT-5.1, and DeepSeek-R1) combined with few-shot prompting and voting strategies. Our experiments show three main findings. First, model diversity and ensemble voting consistently improve performance compared to single-model baselines. Second, the full clinician answer paragraph is provided as additional prompt context for evidence alignment. Third, results on the development set show that alignment accuracy is mainly limited by reasoning. The best scores on the development set reach 88.81 micro F1 on ST4, 65.72 macro F1 on ST2, 34.01 on ST3, and 33.05 on ST1.

电子病历问答系统多模型集成医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。