arXiv:2510.18691cs.CL2025-10

首次研究大模型在长篇医疗问答中的理解能力,揭示了其优劣与改进方向。

Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering

  • 评估不同规模模型在长文本医疗问答中的表现
  • 发现大模型在时间推理任务中仍存局限,RAG在多文档问答中更优
  • 适合关注医疗AI评测与长上下文理解的研究者

本研究首次系统探究大语言模型在长上下文(LC)临床医疗问答任务中的理解能力,突破传统单选题问答(MCQA)范畴。通过涵盖不同内容长度与相关性、多种模型规模及多类任务形式的综合实验,揭示了模型规模的影响及其记忆缺陷,验证了推理型模型的优势,并凸显利用完整患者病历的潜力与挑战。重点考察了检索增强生成(RAG)在单文档与多文档医疗问答中的效果,识别出最佳应用场景。采用多维度评估方法,揭示常见评价指标的局限性。定量分析显示,尽管RAG在多数场景表现优异,但在需要时间推理的任务中仍存在不足。

原文摘要 · Abstract (English)

This study is the first to investigate LLM comprehension capabilities over long-context (LC), clinically relevant medical Question Answering (QA) beyond MCQA. Our comprehensive approach considers a range of settings based on content inclusion of varying size and relevance, LLM models of different capabilities and a variety of datasets across task formulations. We reveal insights on model size effects and their limitations, underlying memorization issues and the benefits of reasoning models, while demonstrating the value and challenges of leveraging the full long patient's context. Importantly, we examine the effect of Retrieval Augmented Generation (RAG) on medical LC comprehension, showcasing best settings in single versus multi-document QA datasets. We shed light into some of the evaluation aspects using a multi-faceted approach uncovering common metric challenges. Our quantitative analysis reveals challenging cases where RAG excels while still showing limitations in cases requiring temporal reasoning.

大模型医疗问答长文本理解RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。