arXiv:2601.06979cs.CL2026-01EMNLP被引 5

用检索增强生成技术,自动从病例报告中提炼教学内容和考题。

MedTutor: A Retrieval-Augmented LLM System for Case-Based Medical Education

  • 结合本地医学教材与最新文献检索,生成精准教育内容。
  • 三位放射科医生评估显示输出具有高临床与教学价值。
  • 适合医学教育者与住院医师提升病例学习效率。

医学住院医师的学习过程面临挑战,需解读复杂病例并快速获取可靠医学知识。传统方式依赖病例研读与师友讨论,但寻找相关教育资源和证据耗时费力。为此,我们提出MedTutor,一种基于检索增强生成(RAG)的系统,可自动从临床病例报告中生成基于证据的教学材料与多选题。该系统采用混合检索机制,协同查询本地医学教材库及学术文献(通过PubMed、Semantic Scholar API),确保内容兼具基础性与前沿性。检索结果经先进重排序模型过滤与排序后,由大语言模型生成最终长文本教学内容。我们进行了严格评估:首先,三位放射科医生评价输出,认为其具备高临床与教育价值;其次,采用大语言模型作为评判者进行大规模评估,结果显示其判断与专家意见存在中等程度相关性,凸显专家监督仍不可或缺。

原文摘要 · Abstract (English)

The learning process for medical residents presents significant challenges, demanding both the ability to interpret complex case reports and the rapid acquisition of accurate medical knowledge from reliable sources. Residents typically study case reports and engage in discussions with peers and mentors, but finding relevant educational materials and evidence to support their learning from these cases is often time-consuming and challenging. To address this, we introduce MedTutor, a novel system designed to augment resident training by automatically generating evidence-based educational content and multiple-choice questions from clinical case reports. MedTutor leverages a Retrieval-Augmented Generation (RAG) pipeline that takes clinical case reports as input and produces targeted educational materials. The system's architecture features a hybrid retrieval mechanism that synergistically queries a local knowledge base of medical textbooks and academic literature (using PubMed, Semantic Scholar APIs) for the latest related research, ensuring the generated content is both foundationally sound and current. The retrieved evidence is filtered and ordered using a state-of-the-art reranking model and then an LLM generates the final long-form output describing the main educational content regarding the case-report. We conduct a rigorous evaluation of the system. First, three radiologists assessed the quality of outputs, finding them to be of high clinical and educational value. Second, we perform a large scale evaluation using an LLM-as-a Judge to understand if LLMs can be used to evaluate the output of the system. Our analysis using correlation between LLMs outputs and human expert judgments reveals a moderate alignment and highlights the continued necessity of expert oversight.

医学教育检索增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。