arXiv:2509.16326cs.CLcs.CV2025-09EMNLP被引 1

为病理报告生成设计了基于实体与关系的评估框架,提升临床质量评判准确性。

HARE: an entity and relation centric evaluation framework for histopathology reports

  • 构建实体与关系为中心的评估体系,聚焦病理关键信息匹配。
  • 在813份病理报告上训练模型,实现91.5%的实体识别准确率。
  • 新指标优于传统方法,更贴近专家评价,适合病理生成研究者使用。

医学领域自动化文本生成是研究热点,但临床质量评估仍具挑战,尤其缺乏特定领域的评估指标,如病理学。本文提出HARE(Histopathology Automated Report Evaluation),一种以实体和关系为核心的新型评估框架,包含基准数据集、命名实体识别(NER)模型、关系抽取(RE)模型及新提出的评估指标。该框架通过对比参考报告与生成报告中关键病理实体与关系的一致性,优先评估临床相关性内容。为构建HARE基准数据集,我们对813份去标识化临床诊断病理报告和652份来自癌症基因组图谱(TCGA)的病理报告进行了领域特异性实体与关系标注。基于领域适配语言模型GatorTronS微调得到的HARE-NER与HARE-RE,在所有测试模型中取得最高整体F1分数(0.915)。所提出的HARE评估指标显著优于传统指标(如ROUGE、Meteor)及放射科评估指标(如RadGraph-XL),在与专家评分的相关性、回归表现方面均领先,优于第二佳方法GREEN(大语言模型驱动的放射科报告评估器):皮尔逊相关系数提升0.168,斯皮曼等级相关系数提升0.161,肯德尔等级相关系数提升0.123,决定系数提高0.176,均方根误差降低0.018。HARE、数据集及模型已开源(https://github.com/knowlab/HARE),旨在推动病理报告生成技术的发展,提供可信赖的评估框架。

原文摘要 · Abstract (English)

Medical domain automated text generation is an active area of research and development; however, evaluating the clinical quality of generated reports remains a challenge, especially in instances where domain-specific metrics are lacking, e.g. histopathology. We propose HARE (Histopathology Automated Report Evaluation), a novel entity and relation centric framework, composed of a benchmark dataset, a named entity recognition (NER) model, a relation extraction (RE) model, and a novel metric, which prioritizes clinically relevant content by aligning critical histopathology entities and relations between reference and generated reports. To develop the HARE benchmark, we annotated 813 de-identified clinical diagnostic histopathology reports and 652 histopathology reports from The Cancer Genome Atlas (TCGA) with domain-specific entities and relations. We fine-tuned GatorTronS, a domain-adapted language model to develop HARE-NER and HARE-RE which achieved the highest overall F1-score (0.915) among the tested models. The proposed HARE metric outperformed traditional metrics including ROUGE and Meteor, as well as radiology metrics such as RadGraph-XL, with the highest correlation and the best regression to expert evaluations (higher than the second best method, GREEN, a large language model based radiology report evaluator, by Pearson $r = 0.168$, Spearman $ρ= 0.161$, Kendall $τ= 0.123$, $R^2 = 0.176$, $RMSE = 0.018$). We release HARE, datasets, and the models at https://github.com/knowlab/HARE to foster advancements in histopathology report generation, providing a robust framework for improving the quality of reports.

病理报告评估框架实体关系临床质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。