arXiv:2503.04176cs.AIcs.CE2025-03被引 45

让大模型学会分析病历的时间线索,提升长期医疗记录推理能力。

TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records

  • 用时间相关的指令-响应对训练模型,强化时序理解
  • 在人工标注和自建基准上分别提升7.3%和9.2%准确率
  • 适合需要处理长期病历的医疗AI研究者和开发者

大型语言模型(LLMs)在医疗任务中展现出巨大潜力,但电子健康记录(EHR)的纵向特性带来了独特挑战。尽管LLMs在医疗任务上的能力持续提升,其对跨多次就诊和时间跨度的时序依赖关系的推理能力仍鲜有探索。本文提出TIMER(Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records),通过将指令-响应对与患者记录的不同时间部分关联,构建时序感知的指令评估与微调框架。我们开发了TIMER-Bench,首个面向纵向EHR的时序推理评估基准;以及TIMER-Instruct,一种用于学习时间推理的指令微调方法。实验表明,经TIMER-Instruct微调的模型在人工生成基准上性能提升7.3%,在TIMER-Bench上提升9.2%,证明时序指令微调能有效增强模型对EHR的时序推理能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have emerged as promising tools for assisting in medical tasks, yet processing Electronic Health Records (EHRs) presents unique challenges due to their longitudinal nature. While LLMs' capabilities to perform medical tasks continue to improve, their ability to reason over temporal dependencies across multiple patient visits and time frames remains unexplored. We introduce TIMER (Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records), a framework that incorporate instruction-response pairs grounding to different parts of a patient's record as a critical dimension in both instruction evaluation and tuning for longitudinal clinical records. We develop TIMER-Bench, the first time-aware benchmark that evaluates temporal reasoning capabilities over longitudinal EHRs, as well as TIMER-Instruct, an instruction-tuning methodology for LLMs to learn reasoning over time. We demonstrate that models fine-tuned with TIMER-Instruct improve performance by 7.3% on human-generated benchmarks and 9.2% on TIMER-Bench, indicating that temporal instruction-tuning improves model performance for reasoning over EHR.

医疗AI时序推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。