arXiv:2605.03618cs.CL2026-05中稿 · CL4Health@LREC 202…

不训练模型,用提示工程评测大模型在医疗问答中的表现。

BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA

  • 用提示工程测试开源与商用大模型在无训练数据下的能力。
  • 最佳组合(MedGemma 3 27B + 好提示)在证据对齐任务中排名第一。
  • 适合医疗低资源场景下无需训练的大模型应用研究者参考。

本文介绍北京理工大学-乌克兰团队(BIT.UA & AAUBS)参与2026年ArchEHR-QA共享任务的成果,聚焦医疗问答与证据定位,在无训练数据且受严格数据隐私限制(如GDPR)的低资源环境下评估大语言模型(LLM)性能。由于无法进行权重更新,我们采用多种提示工程技术,包括任务分解、思维链和上下文学习,对比多个前沿商用模型与可本地部署的开源模型。同时引入多数投票与大模型作为裁判的集成策略以提升预测鲁棒性。结果显示,虽然商用模型对提示变化更具鲁棒性,但经领域适配的开源模型(如MedGemma 3 27B)在搭配合适提示时表现极具竞争力。本方法在子任务4(证据引用对齐)中获第1名,在子任务3(患者友好型回答生成)中获第3名。所有代码、结果与提示均开源于GitHub:https://github.com/bioinformatics-ua/ArchEHR-QA-2026。

原文摘要 · Abstract (English)

This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. Due to the absence of training data and the strict data privacy constraints inherent to the healthcare domain (e.g. GDPR), we investigate the capabilities of Large Language Models (LLMs) without weight updates. We evaluate several state-of-the-art proprietary models and locally deployable open-source alternatives using various prompt engineering strategies, including task decomposition, Chain-of-Thought, and in-context learning. Furthermore, we explore majority voting and LLM-as-a-judge ensembling techniques to maximize predictive robustness. Our results demonstrate that while proprietary models exhibit strong resilience to prompt variations, domain-adapted open-source models (such as MedGemma 3 27B) achieve highly competitive performance when paired with the right prompt. Overall, our prompt-based approach proved highly effective, securing 1st place in Subtask 4 (evidence citation alignment) and 3rd place in Subtask 3 (patient-friendly answer generation). All code, results, and prompts are available on our GitHub repository: https://github.com/bioinformatics-ua/ArchEHR-QA-2026.

医疗问答提示工程开源模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。