arXiv:2508.02082cs.CV2025-08被引 8

构建结构化胸片报告生成数据集与评估框架,提升临床可用性。

S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework

  • 基于大模型生成结构化报告,整合疾病、位置、严重程度等细节。
  • 提出S-Score评估指标,精准衡量关键临床信息的准确性。
  • 适合医学AI研究者及临床辅助系统开发者参考。

放射科报告生成(RRG)在临床和人工智能中具有重要意义。传统自由文本报告存在冗余和语言不一致问题,影响关键临床信息提取。结构化报告生成(S-RRG)通过标准化格式提升可读性,但现有方法依赖预定义标签或模板,输出碎片化且缺乏表达力。本文构建了包含疾病名称、严重程度、概率和解剖位置的胸片结构化数据集(MIMIC-STRUC),训练基于大模型的报告生成器,并提出S-Score评估指标,该指标不仅衡量疾病预测准确率,还评估特定疾病的细节精度,与人工评估高度一致,强调临床决策相关要素。实验表明,结构化报告与定制评估框架能显著提升报告质量。

原文摘要 · Abstract (English)

Radiology report generation (RRG) for diagnostic images, such as chest X-rays, plays a pivotal role in both clinical practice and AI. Traditional free-text reports suffer from redundancy and inconsistent language, complicating the extraction of critical clinical details. Structured radiology report generation (S-RRG) offers a promising solution by organizing information into standardized, concise formats. However, existing approaches often rely on classification or visual question answering (VQA) pipelines that require predefined label sets and produce only fragmented outputs. Template-based approaches, which generate reports by replacing keywords within fixed sentence patterns, further compromise expressiveness and often omit clinically important details. In this work, we present a novel approach to S-RRG that includes dataset construction, model training, and the introduction of a new evaluation framework. We first create a robust chest X-ray dataset (MIMIC-STRUC) that includes disease names, severity levels, probabilities, and anatomical locations, ensuring that the dataset is both clinically relevant and well-structured. We train an LLM-based model to generate standardized, high-quality reports. To assess the generated reports, we propose a specialized evaluation metric (S-Score) that not only measures disease prediction accuracy but also evaluates the precision of disease-specific details, thus offering a clinically meaningful metric for report quality that focuses on elements critical to clinical decision-making and demonstrates a stronger alignment with human assessments. Our approach highlights the effectiveness of structured reports and the importance of a tailored evaluation metric for S-RRG, providing a more clinically relevant measure of report quality.

结构化报告医学AI评估指标大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。