arXiv:2601.03018cs.CLcs.AI2026-01被引 1

用强化学习从病历中推理痴呆发展轨迹,提升真实世界预测效果。

Dementia-R1: Reinforced Pretraining and Reasoning from Unstructured Clinical Notes for Real-World Dementia Prognosis

  • 先预训练模型预测可验证的临床指标,再通过强化学习推理疾病进展。
  • 在真实数据集上达到84.02%的AUROC,优于大10倍的模型。
  • 适用于痴呆、帕金森痴呆等长期病情预测,适合临床研究者使用。

尽管大型语言模型在临床文本理解上表现优异,但在痴呆预判这类纵向预测任务中仍面临挑战,因其需对多轮就诊中复杂且非单调的症状演变进行推理。标准监督训练缺乏症状演化的显式标注,而直接强化学习受限于稀疏的二值奖励。为此,我们提出Dementia-R1,一种基于强化学习的纵向痴呆预判框架,利用未结构化病历实现推理。该方法采用冷启动强化学习策略,先让模型预训练以预测从患者病史中提取的可验证临床指标,从而增强其对疾病进程的推理能力,再决定最终临床状态。大量实验表明,Dementia-R1在AMC真实世界未结构化队列中表现最佳,达到84.02%的AUROC,优于最大达10倍的模型。该框架在独立医院队列中也成功应用于帕金森病痴呆预测,取得78.37%的AUROC。在ADNI基准上,我们的7B模型在所有LLM基线中取得最高AUROC(83.17%),展现出对波动性认知轨迹的强大纵向推理能力。代码已公开于https://anonymous.4open.science/r/dementiar1-CDB5。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have shown strong performance on clinical text understanding, they struggle with longitudinal prediction tasks such as dementia prognosis, which require reasoning over complex, non-monotonic symptom trajectories across multiple visits. Standard supervised training lacks explicit annotations for symptom evolution, while direct Reinforcement Learning (RL) is hindered by sparse binary rewards. To address this challenge, we introduce Dementia-R1, an RL-based framework for longitudinal dementia prognosis from unstructured clinical notes. Our approach adopts a Cold-Start RL strategy that pre-trains the model to predict verifiable clinical indices extracted from patient histories, enhancing the capability to reason about disease progression before determining the final clinical status. Extensive experiments show that Dementia-R1 achieves the best overall performance on the AMC real-world unstructured cohort, reaching an AUROC of 84.02% and outperforming models up to 10x larger. The framework also generalizes to Parkinson's disease dementia prediction in an independent hospital cohort, achieving an AUROC of 78.37%. On the ADNI benchmark, our 7B model attains the highest AUROC among all LLM baselines at 83.17%, demonstrating strong longitudinal reasoning over fluctuating cognitive trajectories. Code is available at https://anonymous.4open.science/r/dementiar1-CDB5.

痴呆预测强化学习纵向分析临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。