对比医生与AI在医学问答中对信息重要性的判断差异。
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
- 构建医患问答数据集,标注每句话的相关性。
- 发现模型与医生对关键信息的判断常不一致。
- 过滤无关内容后,医生和模型准确率均提升。
大型语言模型(LLMs)在多项医学问答基准测试中表现优异,包括标准化医学考试。然而,正确答案并不保证推理过程正确,模型可能通过错误逻辑得出正确结论。本研究提出MedPAIR(医学问答中医生与AI相关性评估数据集),用于评估医师训练者与LLM在回答医学问题时对相关信息的优先级判断。我们从36名医师训练者处获取1,300个问答对的标注,对问题组件中的每句话进行相关性标记。将这些相关性估计与LLM的结果进行比较,并进一步评估“相关”信息子集对下游任务性能的影响。结果表明,LLM与医师训练者的内容相关性判断存在显著偏差。在剔除医师标记的无关句子后,医师与LLM的准确率均有提升。所有标注数据及模型结果均可在http://medpair.csail.mit.edu/获取。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may reach accurate conclusions through flawed processes. In this study, we introduce the MedPAIR (Medical Dataset Comparing Physicians and AI Relevance Estimation and Question Answering) dataset to evaluate how physician trainees and LLMs prioritize relevant information when answering QA questions. We obtain annotations on 1,300 QA pairs from 36 physician trainees, labeling each sentence within the question components for relevance. We compare these relevance estimates to those for LLMs, and further evaluate the impact of these "relevant" subsets on downstream task performance for both physician trainees and LLMs. We find that LLMs are frequently not aligned with the content relevance estimates of physician trainees. After filtering out physician trainee-labeled irrelevant sentences, accuracy improves for both the trainees and the LLMs. All LLM and physician trainee-labeled data are available at: http://medpair.csail.mit.edu/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。