arXiv:2601.16753cs.CLcs.AI2026-01

用大模型自动标注放射科报告中的纵向变化,提升评估一致性。

Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation

  • 用大模型自动识别报告中随时间变化的病灶信息
  • 在500份人工标注报告上验证,性能比现有方法高5.3%~11.3%
  • 生成超9万份标注数据,可作新基准供模型评估

放射科报告中的纵向信息指多次检查中病灶的动态变化,对监测病情和临床决策至关重要。尽管近年许多自动化报告生成方法试图捕捉此类信息,但其性能验证仍缺乏统一标注工具。本文提出一种基于大语言模型(LLM)的自动标注流水线,先识别相关句子,再提取疾病进展。我们在500份人工标注报告上评估五种主流LLM,最终选用Qwen2.5-32B对MIMIC-CXR公开数据集中的95,169份报告进行标注。该标注数据集构建了标准化评估基准,用于评测七种先进报告生成模型。实验表明,本方法在纵向信息检测与疾病追踪任务上分别取得11.3%和5.3%更高的F1-score,显著优于现有方案。源代码已开源。

原文摘要 · Abstract (English)

Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is crucial for monitoring disease progression and guiding clinical decisions. Many recent automated radiology report generation methods are designed to capture longitudinal information; however, validating their performance is challenging. There is no proper tool to consistently label temporal changes in both ground-truth and model-generated texts for meaningful comparisons. Large language models (LLMs) offer a promising annotation alternative, as they are capable of capturing nuanced linguistic patterns and semantic similarities without extensive manual intervention. They also adapt well to new contexts. In this study, we therefore propose an LLM-based pipeline to automatically annotate longitudinal information in radiology reports. The pipeline first identifies sentences containing relevant information and then extracts the progression of diseases. We evaluate and compare five mainstream LLMs on these two tasks using 500 manually annotated reports. Considering both efficiency and performance, Qwen2.5-32B was subsequently selected and used to annotate another 95,169 reports from the public MIMIC-CXR dataset. Our Qwen2.5-32B-annotated dataset provided us with a standardized benchmark for evaluating report generation models. Using this new benchmark, we assessed seven state-of-the-art report generation models. Our LLM-based annotation method outperforms existing annotation solutions, achieving 11.3\% and 5.3\% higher F1-scores for longitudinal information detection and disease tracking, respectively. The source code is available at https://github.com/wxinyi1996/Standardizing-Longitudinal-Chest-X-ray-Report-Evaluation-via-Large-Language-Model-Annotation.git.

医学报告大模型纵向分析标注标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。