用视觉语言模型+低秩微调,小数据下精准定位手术工具2D关键点。
Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation
- 用提示词对齐视觉特征与语义描述,构建指令微调数据集。
- 仅两轮微调即超越基线模型,验证低秩适配在小样本中的有效性。
- 适合医疗视觉任务研究者,为3D姿态估计提供新思路。
本文提出一种新颖的2D手术工具关键点检测流程,利用经过低秩适配(LoRA)微调的视觉语言模型(VLMs)。相较于传统卷积神经网络(CNN)或基于Transformer的方法在小规模医学数据集上易过拟合的问题,本方法借助预训练VLM的泛化能力。通过精心设计提示词构建指令微调数据集,实现视觉特征与语义关键点描述的对齐。实验表明,仅需两轮微调,适配后的VLM即超越基线模型,证明了LoRA在低资源场景下的有效性。该方法不仅提升了关键点检测性能,也为未来3D手术手及工具姿态估计研究奠定基础。
原文摘要 · Abstract (English)
This paper presents a novel pipeline for 2D keypoint estima- tion of surgical tools by leveraging Vision Language Models (VLMs) fine- tuned using a low rank adjusting (LoRA) technique. Unlike traditional Convolutional Neural Network (CNN) or Transformer-based approaches, which often suffer from overfitting in small-scale medical datasets, our method harnesses the generalization capabilities of pre-trained VLMs. We carefully design prompts to create an instruction-tuning dataset and use them to align visual features with semantic keypoint descriptions. Experimental results show that with only two epochs of fine tuning, the adapted VLM outperforms the baseline models, demonstrating the ef- fectiveness of LoRA in low-resource scenarios. This approach not only improves keypoint detection performance, but also paves the way for future work in 3D surgical hands and tools pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。