arXiv:2608.17519cs.CVcs.LG2026-08

研究手术技能模型能否跨评分体系迁移,发现视觉特征主导但非唯一决定因素。

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

论文配图:Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?
图 1 · 摘自论文原文
  • 通过对比不同训练方法,检验模型在不同评分体系间的迁移能力。
  • JIGSAWS预训练模型在LASANA上表现良好(CCC 0.77-0.80),但反向迁移失败。
  • 提示任务特定头是预测关键,骨干网络只需提供时空特征即可。

基于视觉的手术技能评估在域内表现优异,但一个根本问题尚未被探讨:这些模型是否学习到可迁移的手术熟练度表示,还是仅编码了数据集特有的视觉模式?本文系统分析了在LASANA和JIGSAWS数据集上,使用GOALS与OSATS评估标准时跨域技能迁移的限制。每种方法均有特定诊断目的:端到端训练用于测试监督学习是否可直接迁移,自适应尖锐感知最小化(ASAM)用于探究更平坦的损失景观是否提升泛化能力,基于增强的自监督与对比学习用于评估领域不变预训练是否能解耦技能与视觉上下文。采用不重叠参与者留出测试集评估双向迁移。结果显示不对称性:在目标域提供一致监督的情况下,基于JIGSAWS预训练的主干网络在LASANA上达到CCC值0.77至0.80,接近端到端基线,表明跨评分体系迁移可行;但在所有方法下向JIGSAWS迁移均失败,可能源于标注不一致。控制实验显示,使用Kinetics预训练主干网络时,任务特定头承担了大部分技能预测责任,而主干网络只需提供充分的时空特征。这些发现揭示了一个新视角:此前未被研究的核心问题是技能表示能否在不同评分系统间迁移。结果表明视觉成分占主导地位,但并非唯一决定因素;需进一步工作以明确可迁移技能特征与绑定特定视觉域的特征之间的区别。

原文摘要 · Abstract (English)

Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment scales using the LASANA and JIGSAWS datasets. Each evaluated method serves a targeted diagnostic purpose: end-to-end training to test whether supervised skill learning transfers directly, Adaptive Sharpness-Aware Minimization (ASAM) to probe whether flatter loss landscapes improve generalization, and augmentation-based self-supervised and contrastive learning to assess whether domain-invariant pretraining decouples skill from visual context. Transfer is evaluated in both directions using a disjoint-participant held-out test set for JIGSAWS. Results reveal an asymmetry: backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching the end-to-end baseline, showing cross-rubric transfer is feasible when the target domain provides consistent supervision. Transfer to JIGSAWS fails across all methods, likely due to annotation inconsistencies. Control experiments with a Kinetics-pretrained backbone suggest task-specific heads carry the majority of the skill prediction burden, while the backbone need only provide adequate spatiotemporal features. These findings offer a new perspective on vision-based skill assessment: the central question of whether skill representations transfer across scoring systems has not been previously investigated. Results indicate the visual component is dominant but not solely responsible for skill prediction; further work is needed to conclusively disentangle transferable skill features from those bound to a specific visual domain.

手术评估模型迁移视觉表征多尺度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。