arXiv:2510.27158cs.CV2025-10

用AI评估肾移植活检的巴夫评分,发现现有模型仍难准确还原专家判断。

How Close Are We? Limitations and Progress of AI Models in Banff Lesion Scoring

  • 将巴夫评分拆解为结构与炎症组件,用规则框架整合检测结果
  • 部分指标可匹配专家评分,但中间过程常出现漏检或误判
  • 强调需模块化评估,推动病理AI标准化发展

巴夫分类是全球肾移植活检评估的标准,但其半定量特性、复杂标准及观察者间差异给计算复现带来挑战。本研究通过模块化规则框架,探索现有深度学习模型近似巴夫病变评分的可行性。我们将每项指标(如肾小球炎g、管周毛细血管炎ptc、内膜动脉炎v)分解为结构与炎症成分,评估当前分割与检测工具是否支持其计算。模型输出依据与专家指南一致的启发式规则映射为巴夫评分,并与专家标注真值对比。结果揭示部分成功与关键失败模式:结构遗漏、幻觉生成和检测模糊。即使最终评分匹配专家标注,中间表示不一致也削弱可解释性。研究揭示当前AI流程在复制专家级分级方面的局限,强调模块化评估与计算巴夫评分标准对推动移植病理AI发展的必要性。

原文摘要 · Abstract (English)

The Banff Classification provides the global standard for evaluating renal transplant biopsies, yet its semi-quantitative nature, complex criteria, and inter-observer variability present significant challenges for computational replication. In this study, we explore the feasibility of approximating Banff lesion scores using existing deep learning models through a modular, rule-based framework. We decompose each Banff indicator - such as glomerulitis (g), peritubular capillaritis (ptc), and intimal arteritis (v) - into its constituent structural and inflammatory components, and assess whether current segmentation and detection tools can support their computation. Model outputs are mapped to Banff scores using heuristic rules aligned with expert guidelines, and evaluated against expert-annotated ground truths. Our findings highlight both partial successes and critical failure modes, including structural omission, hallucination, and detection ambiguity. Even when final scores match expert annotations, inconsistencies in intermediate representations often undermine interpretability. These results reveal the limitations of current AI pipelines in replicating computational expert-level grading, and emphasize the importance of modular evaluation and computational Banff grading standard in guiding future model development for transplant pathology.

医学图像分析肾移植巴夫评分AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。