构建标准化评测体系,让病理报告生成更真实可信。
PathReportEval: A Systematic Benchmark for Pathology Report Generation

- 统一数据、模型与评估流程,支持公平对比。
- 提出临床质量评分CRQS,精准捕捉诊断错误。
- 适合医学AI研究者与临床交叉团队使用。
从全幻灯片图像(WSIs)生成病理报告是快速发展的多模态学习任务,但进展难以衡量,因现有研究使用异构数据集、模型设置、视觉编码器和评估协议。常用自然语言生成指标(如BLEU、ROUGE、METEOR)主要奖励词汇相似性,常无法检测遗漏诊断、虚构发现或肿瘤特征矛盾等临床关键错误。本文提出标准化基准与评估框架,评估四种代表性方法在三个数据集(TCGA、HistAI、REG 2025)上表现,使用三种病理基础编码器(CONCHv1.5、UNI2-h、H-Optimus-1)。框架统一预处理、特征提取、训练、解码与评估流程,支持新方法、数据集与编码器的模块化集成。核心贡献为临床报告质量评分(CRQS),基于结构化临床属性评估事实正确性,涵盖四个维度:临床事实覆盖度、关键信息召回率、幻觉率与临床矛盾度,提供总分与可解释子分。实验表明,传统语言指标与临床正确性弱相关,常高估报告质量;而CRQS揭示了模型与编码器间临床意义差异,传统指标未捕捉到。该基准、开源框架与CRQS共同建立可复现的病理报告生成评估基础。
原文摘要 · Abstract (English)
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders, and evaluation protocols. Moreover, commonly used natural language generation metrics, including BLEU, ROUGE, and METEOR, primarily reward lexical similarity and often fail to detect clinically consequential errors such as omitted diagnoses, hallucinated findings, or discordant tumor attributes. We present a standardized benchmark and evaluation framework for pathology report generation. The benchmark evaluates four representative methods across three datasets (TCGA, HistAI, and REG 2025) using three pathology foundation encoders (CONCHv1.5, UNI2-h, and H-Optimus-1). Our framework standardizes preprocessing, feature extraction, training, decoding, and evaluation, enabling fair comparison across models while providing a modular platform for integrating new methods, datasets, and encoders. A central contribution is the Clinical Report Quality Score (CRQS), a clinically grounded metric for evaluating factual correctness. CRQS maps reference and generated reports into structured clinical attributes and measures four complementary dimensions: clinical fact coverage, key information recall, hallucination rate, and clinical discordance, producing both an overall score and interpretable sub-scores. Experiments demonstrate that conventional language-generation metrics are weakly aligned with clinical correctness and frequently overestimate report quality. In contrast, CRQS reveals clinically meaningful differences between models and encoders that lexical metrics fail to capture. Together, the benchmark, public plug-and-play framework, and CRQS establish a reproducible foundation for rigorous evaluation of pathology report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。