用大模型迭代生成并修正临床症状时间线,提升病程推理准确性。
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

- 基于大模型生成与约束验证的双组件框架,逐步优化症状时间序列。
- 在5347条疫苗不良反应文本上,3166份报告实现专家验证的时间阶段标注。
- 适用于医学文本时间推理,尤其适合缺乏明确时间锚点的病历分析。
理解临床叙述中症状的时间演变对疾病监测、安全监控和因果关系评估至关重要。然而,临床叙述通常缺乏明确的时间锚点。现有方法主要聚焦于多访视、带时间戳记录中的成对关系分类,而对单个时间锚点稀疏报告中结构化症状轨迹的重建仍缺乏有效手段。我们提出CRAFT,一个基于大模型的框架,结合生成器与基于约束的验证器,通过定向反馈实现分阶段症状时间线的迭代生成与优化。我们在MedTempo上进行评估,该基准包含5,347条涵盖三种新冠疫苗的不良事件叙述,其中3,166份报告经专家验证具有时间阶段标注。在四种大模型基座上的实验表明,CRAFT在时间顺序准确性上持续提升;消融分析进一步揭示了生成器与验证器组件在不同模型能力水平下的贡献。
原文摘要 · Abstract (English)
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。