用轻量神经网络提升脂肪肝纤维化无创检测,效果优于传统公式和大模型。
Machine-Learning-Enhanced Non-Invasive Testing for MASLD Fibrosis: Shallow-Deep Neural Networks Versus FIB-4, Tabular Foundation Models, and Large Language Models

- 用浅层深度网络融合五项临床指标,优化纤维化风险预测。
- 在马来西亚和印度外数据集上,新模型AUC达0.77和0.67,超越FIB-4。
- 模型仅354个参数却表现稳健,适合临床部署,尤其关注成本与效率者必看。
代谢功能障碍相关脂肪性肝病(MASLD)中,肝纤维化是导致肝病发病率的关键因素。目前广泛使用的无创检测工具FIB-4因固定公式可能未充分利用年龄、天冬氨酸氨基转移酶(AST)、丙氨酸氨基转移酶(ALT)及血小板计数等信息。本研究评估机器学习增强型无创检测(MLE-NIT)能否在保留原FIB-4变量空间的前提下提升高级纤维化检测能力。纳入中国、马来西亚和印度三个经活检证实的MASLD队列(n=784),其中中国队列分为486例训练集与54例内部验证/调参集;最终性能仅在马莱西亚和印度两个外部队列报告。模型使用五项变量:年龄、FIB-4、AST、ALT、血小板计数。对比了FIB-4、浅层-深度神经网络(s-DNN)、TabPFN和gpt-4o-2024-08-06。FIB-4在马来西亚和印度队列的外部ROC-AUC分别为0.75和0.60;TabPFN为0.69和0.66;微调后的GPT-4o为0.75和0.63;s-DNN则分别达到0.77和0.67。尽管s-DNN仅有354个可训练参数(远低于TabPFN的7,244,554),但其外部表现更均衡。校准结果显示s-DNN的布里尔分数为0.18和0.22,置换重要性分析表明AST和FIB-4为关键变量。紧凑的非线性机器学习模型可在不增加临床数据需求的情况下,有效提升基于FIB-4的纤维化评估水平。
原文摘要 · Abstract (English)
Advanced fibrosis is a major determinant of liver-related morbidity in metabolic dysfunction-associated steatotic liver disease (MASLD). FIB-4 is widely used as a first-line non-invasive test, but its fixed formula may underuse diagnostic information contained in age, aspartate aminotransferase, alanine aminotransferase, and platelet count. We evaluated whether machine-learning-enhanced non-invasive testing (MLE-NIT) can improve advanced fibrosis detection while preserving this FIB-4 variable space. We used three biopsy-confirmed MASLD cohorts from China, Malaysia, and India (n=784). The Chinese cohort was split into 486 training and 54 internal validation/tuning patients; final performance was reported only on the Malaysian and Indian external cohorts. Models used five variables: age, FIB-4, aspartate aminotransferase, platelet count, and alanine aminotransferase. We compared FIB-4 with a shallow-deep neural network (s-DNN), TabPFN, and gpt-4o-2024-08-06. FIB-4 achieved external ROC-AUCs of 0.75 and 0.60 in Malaysia and India, respectively. TabPFN achieved 0.69 and 0.66, fine-tuned GPT-4o achieved 0.75 and 0.63, and the s-DNN achieved 0.77 and 0.67, respectively. The s-DNN contained only 354 trainable parameters, compared with 7,244,554 for TabPFN, yet provided a more balanced external operating profile. Calibration showed s-DNN Brier scores of 0.18 and 0.22, and permutation importance identified AST and FIB-4 as dominant variables. Compact non-linear MLE-NITs may enhance FIB-4-based fibrosis assessment without increasing clinical data requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。