提出答案优先的临床问答方法,提升证据定位与对齐效果
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering

- 先生成候选答案并引用原文句子,再判断证据相关性
- 证据识别严格微F1达62.90,答案-证据对齐F1为79.81
- 适合关注临床NLP可读性与系统可解释性的研究者
我们介绍UIC-AIHealth4All系统参与ArchEHR-QA 2026共享任务,涵盖证据识别、答案生成与答案-证据对齐三个子任务。在子任务2和3中,提出答案优先的流水线:模型先生成带具体句引用的候选答案,再分类完整证据集,利用摘要相关性判断与生成答案之间的不对称性。在子任务4中,通过五次独立模型调用的自一致性投票,按阈值保留链接。该方案在证据识别中排名第三(严格微F1 62.90),答案生成第九(总体得分31.90),答案-证据对齐第五(F1 79.81)。对45个语言特征的后验分析显示,模型输出虽词数和句数与临床参考一致,但弗莱施-金凯德等级仍高3.2级,表明临床NLP系统需显式优化可读性。代码与提示语见https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all。
原文摘要 · Abstract (English)
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self-consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer-evidence alignment (F1 79.81). A post-hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch-Kincaid grade levels harder to read than clinician-authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。