arXiv:2511.04255cs.CVcs.AI2025-11被引 1

用人体姿态模型做医学影像关键点检测,性能显著提升

MedSapiens: Taking a Pose to Rethink Medical Imaging Landmark Detection

  • 用多数据集预训练将人体姿态模型Sapiens迁移到医学影像
  • 在多个数据集上平均成功检测率提升21.81%(相比专用模型)
  • 适合少样本场景,少量标注下仍优于现有方法

本文不提出新架构,而是重新审视一个被忽视的基础方案:将面向人体姿态估计的通用视觉基础模型Sapiens适配到医学影像关键点检测任务。尽管传统方法依赖领域特定模型,但大规模预训练视觉模型带来了新可能。本研究通过多数据集预训练,将Sapiens模型适配至医学影像,建立新基准。所提出的MedSapiens模型表明,本已优化空间姿态定位的人体中心基础模型,为解剖学关键点检测提供了强大先验,但该潜力长期未被挖掘。我们在多个数据集上对比现有最优模型,实现平均成功检测率(SDR)提升最高达5.26%(相对于通用模型)和21.81%(相对于专用模型)。进一步评估其在低标注数据下的适应性,在少样本设置中较当前少样本最优表现提升2.69%。代码与模型权重已开源。

原文摘要 · Abstract (English)

This paper does not introduce a novel architecture; instead, it revisits a fundamental yet overlooked baseline: adapting human-centric foundation models for anatomical landmark detection in medical imaging. While landmark detection has traditionally relied on domain-specific models, the emergence of large-scale pre-trained vision models presents new opportunities. In this study, we investigate the adaptation of Sapiens, a human-centric foundation model designed for pose estimation, to medical imaging through multi-dataset pretraining, establishing a new state of the art across multiple datasets. Our proposed model, MedSapiens, demonstrates that human-centric foundation models, inherently optimized for spatial pose localization, provide strong priors for anatomical landmark detection, yet this potential has remained largely untapped. We benchmark MedSapiens against existing state-of-the-art models, achieving up to 5.26% improvement over generalist models and up to 21.81% improvement over specialist models in the average success detection rate (SDR). To further assess MedSapiens adaptability to novel downstream tasks with few annotations, we evaluate its performance in limited-data settings, achieving 2.69% improvement over the few-shot state of the art in SDR. Code and model weights are available at https://github.com/xmed-lab/MedSapiens .

医学影像关键点检测基础模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。