首个真实世界肝纤维化分期AI评测基准,揭示当前水平与临床应用差距。
How Far Has AI Come in Liver Fibrosis Staging? A Large-Scale Real-World Dataset and Benchmark

- 基于多中心MRI数据构建大规模真实世界评测集
- 顶尖AI模型在部分场景已达资深医生水平
- 跨中心差异和数据不平衡是主要挑战,适合医学AI研究者参考
尽管方法持续进步,但人工智能在肝纤维化分期中的实际进展始终缺乏在异质、多中心临床环境下系统的评估。为此,我们引入LiFS——一个源自MICCAI 2025 CARE-Liver挑战的大规模数据集与基准,包含610名患者,来自多个中心与扫描仪的多序列MRI数据。据我们所知,LiFS是首个提供完整钆塞酸增强序列并经组织病理学验证标注的真实世界扫描仪多样性数据集。通过系统评估9个独立开发的方法(从96支参赛团队中选出),对比队内放射科医生参考结果,从三个互补视角揭示当前AI在肝纤维化分期上的进展:首先,在特定场景下,最优AI模型与资深放射科医生相当,显著优于初级医生;中位数水平则接近初级医生。其次,数据层面显示,跨中心异质性、标签不平衡及增强序列差异是主要挑战。第三,技术层面表明,空间配准、输入维度、多模态融合策略与主干网络架构等设计选择影响跨中心鲁棒性,但单一选择无法完全弥合差距。总体而言,LiFS为定位当前肝纤维化分期AI状态提供了严谨的现实基准,并推动未来研究解决临床可靠部署的关键瓶颈。
原文摘要 · Abstract (English)
Despite years of methodological progress, how far AI has come in liver fibrosis staging has never been systematically evaluated under the heterogeneous, multi-center conditions that define clinical practice. To address this gap, we introduce LiFS, a large-scale dataset and benchmark derived from the MICCAI 2025 CARE-Liver challenge, comprising 610 patients across multiple centers and scanners with multi-sequence MRI. To the best of our knowledge, LiFS is the first benchmark providing complete gadoxetic acid-enhanced sequences with histopathology-confirmed annotations from diverse real-world scanners. Through systematic evaluation of 9 independently developed methods selected from 96 registered teams against in-cohort radiologist reference results, our findings address how far current AI has progressed toward clinical-level liver fibrosis staging from three complementary perspectives. First, against radiologists, the best AI methods were broadly comparable to the senior radiologist and significantly exceeded the junior radiologist in selected settings, while median AI performance generally approached junior-radiologist levels. Second, from a data perspective, cross-center heterogeneity, label imbalance, and contrast-enhanced sequence variability emerge as the dominant challenges for AI methods. Third, from a technical perspective, methodological design choices, including spatial registration, input dimensionality, multi-modal fusion strategy, and backbone architecture, appear to modulate cross-center robustness, although no single choice alone closes the gap. Overall, LiFS provides a rigorous real-world benchmark for positioning the current state of AI in liver fibrosis staging and for enabling future research on the key challenges that limit clinically reliable deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。