大语言模型在胎心监测分析中表现超越专用模型,但需权衡计算成本。
Large language models surpass domain-specific architectures for antepartum electronic fetal monitoring analysis
- 用统一框架对比15种模型,测试2500段胎心监测数据
- 微调后的大语言模型在多数情况下准确率最高,尤其有宫缩信号时
- 适合追求高精度的医疗研究者,但需考虑算力限制
基础模型(FMs)和大语言模型(LLMs)在时间序列分析中展现出良好泛化能力,但在电子胎儿监测(EFM)和胎心图(CTG)分析中的潜力尚未充分探索。现有大多数CTG研究依赖领域专用模型,且缺乏与现代基础模型或语言模型的系统性比较,难以判断这些模型是否能在胎儿健康评估中超越专用系统。本研究首次对先进架构在产前CTG分类任务中进行综合性基准测试。基于超过2500段20分钟的记录,在统一框架下评估了15种涵盖领域专用、时间序列、基础模型及语言模型的架构。微调后的语言模型在不同数据可用性场景下均表现优于基础模型和领域专用模型,仅在无宫缩信号时,领域专用模型更具鲁棒性。然而,性能提升需付出更高计算资源代价。结果表明,尽管微调语言模型在CTG分类上达到当前最优水平,实际部署仍需在性能与计算效率间权衡。
原文摘要 · Abstract (English)
Foundation models (FMs) and large language models (LLMs) have demonstrated promising generalization across diverse domains for time-series analysis, yet their potential for electronic fetal monitoring (EFM) and cardiotocography (CTG) analysis remains underexplored. Most existing CTG studies relied on domain-specific models and lack systematic comparisons with modern foundation or language models, limiting our understanding of whether these models can outperform specialized systems in fetal health assessment. In this study, we present the first comprehensive benchmark of state-of-the-art architectures for automated antepartum CTG classification. Over 2,500 20-minutes recordings were used to evaluate over 15 models spanning domain-specific, time-series, foundation, and language-model categories under a unified framework. Fine-tuned LLMs consistently outperformed both foundation and domain-specific models across data-availability scenarios, except when uterine-activity signals were absent, where domain-specific models showed greater robustness. These performance gains, however, required substantially higher computational resources. Our results highlight that while fine-tuned LLMs achieved state-of-the-art performance for CTG classification, practical deployment must balance performance with computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。