为CT报告生成设计细粒度诊断评估基准,精准检测临床错误。
CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation

- 基于问答对构建细粒度临床属性评估体系
- 在专家评估中相关性更高,误检率显著降低
- 适合临床可信度验证与模型精细调优
CT报告生成的评估仍面临重大挑战,源于文本量大、病灶多样复杂,以及疾病导向的细粒度属性。传统评估指标仅提供粗粒度的词汇重叠或实体匹配,无法反映临床所需的细粒度诊断准确性。为此,我们提出CT-FineBench,一个基于CT-RATE和Merlin构建的细粒度事实一致性评估基准。该基准通过问答(QA)流程构建:首先识别并结构化关键病灶特异性临床属性(如位置、大小、边界);其次系统地将这些属性转化为问答数据集,问题针对金标准报告中的具体临床细节。评估协议使用该问答集查询机器生成报告,并评分答案正确性。该方法实现全面、可解释、临床相关的评估,超越表面词汇重叠,精准定位具体临床错误。实验表明,CT-FineBench与专家临床评估的相关性更强,对细粒度事实错误的敏感度远超以往指标。
原文摘要 · Abstract (English)
The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of findings, and the presence of fine-grained, disease-oriented attributes. Conventional evaluation metrics offer only coarse measures of lexical overlap or entity matching and fail to reflect the granular diagnostic accuracy required for clinical use. To address this gap, we propose CT-FineBench, a benchmark built from CT-RATE and Merlin to evaluate the fine-grained factual consistency of CT reports, constructed from CT-RATE and Merlin. Our benchmark is constructed through a meticulous, Question-Answering (QA) based process: first, we identify and structure key, finding-specific clinical attributes (like location, size, margin). Second, we systematically transform these attributes into a QA dataset, where questions probe for specific clinical details grounded in gold-standard reports. The evaluation protocol for CT-FineBench involves using this QA dataset to query a machine-generated report and scoring the correctness of the answers. This allows for a comprehensive, interpretable, and clinically-relevant assessment, moving beyond superficial lexical overlap to pinpoint specific clinical errors. Experiments show that CT-FineBench correlates better with expert clinical assessment and is substantially more sensitive to fine-grained factual errors than prior metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。