arXiv:2508.16730eess.IVcs.CV2025-08中稿 · DEMI workshop MICC…被引 2

评估三种模型迁移能力指标,帮医生选最适合的手术阶段识别模型。

Analysis of Transferability Estimation Metrics for Surgical Phase Recognition

  • 用嵌入特征预测模型在新数据上的微调效果,无需重训练。
  • LogME指标最准,尤其最小子集得分组合表现最佳。
  • 适合医疗影像领域研究者快速筛选预训练模型。

微调预训练模型已成为现代机器学习的核心方法,使在标注数据有限的情况下仍能获得高性能。在手术视频分析中,专家标注耗时且成本高,因此选择最适合下游任务的预训练模型既关键又具挑战性。源无关迁移能力估计(SITE)通过仅使用模型的嵌入或输出,预测其在目标数据上微调的效果,无需完整重训练。本文首次将SITE形式化应用于手术阶段识别,并在两个不同数据集(RAMIE和AutoLaparo)上对三种代表性指标(LogME、H-Score、TransRate)进行综合基准测试。结果表明,LogME(尤其是按子集最小得分聚合)与微调准确率最接近;H-Score预测能力弱;TransRate常出现排名反转。消融实验显示,当候选模型性能相近时,迁移能力估计的区分力下降,强调保持模型多样性或使用额外验证的重要性。文章最后给出实用的模型选择建议,并展望面向特定领域的度量、理论基础及交互式基准工具的发展方向。

原文摘要 · Abstract (English)

Fine-tuning pre-trained models has become a cornerstone of modern machine learning, allowing practitioners to achieve high performance with limited labeled data. In surgical video analysis, where expert annotations are especially time-consuming and costly, identifying the most suitable pre-trained model for a downstream task is both critical and challenging. Source-independent transferability estimation (SITE) offers a solution by predicting how well a model will fine-tune on target data using only its embeddings or outputs, without requiring full retraining. In this work, we formalize SITE for surgical phase recognition and provide the first comprehensive benchmark of three representative metrics, LogME, H-Score, and TransRate, on two diverse datasets (RAMIE and AutoLaparo). Our results show that LogME, particularly when aggregated by the minimum per-subset score, aligns most closely with fine-tuning accuracy; H-Score yields only weak predictive power; and TransRate often inverses true model rankings. Ablation studies show that when candidate models have similar performances, transferability estimates lose discriminative power, emphasizing the importance of maintaining model diversity or using additional validation. We conclude with practical guidelines for model selection and outline future directions toward domain-specific metrics, theoretical foundations, and interactive benchmarking tools.

手术识别迁移学习模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。