arXiv:2602.12002cs.CV2026-02被引 1

用微调的局部视觉语言模型提升新生儿复苏动作识别准确率

Can Local Vision-Language Models improve Activity Recognition over Vision Transformers? -- Case Study on Newborn Resuscitation

  • 用小规模视觉语言模型结合LoRA微调,提升细粒度动作识别
  • 微调后模型F1达0.91,优于TimeSFormer的0.70
  • 适合临床视频分析与生成式AI医疗应用研究者

新生儿复苏的精准记录对质量改进和遵循临床指南至关重要,但目前仍未能充分应用。先前使用3D-CNN和视觉变压器(ViT)的方法在从新生儿复苏视频中检测关键动作方面表现良好,但仍面临细粒度动作识别的挑战。本文探究生成式AI方法在提升此类视频动作识别中的潜力,具体研究局部视觉语言模型(VLM)与大语言模型(LLM)结合的效果,并与监督式TimeSFormer基线进行对比。基于包含13.26小时新生儿复苏视频的模拟数据集,评估了多种零样本VLM策略及带分类头的微调VLM方法,包括低秩适应(LoRA)。结果表明,小型(局部)VLM存在幻觉问题,但经LoRA微调后,F1得分达到0.91,超过TimeSformer的0.70。

原文摘要 · Abstract (English)

Accurate documentation of newborn resuscitation is essential for quality improvement and adherence to clinical guidelines, yet remains underutilized in practice. Previous work using 3D-CNNs and Vision Transformers (ViT) has shown promising results in detecting key activities from newborn resuscitation videos, but also highlighted the challenges in recognizing such fine-grained activities. This work investigates the potential of generative AI (GenAI) methods to improve activity recognition from such videos. Specifically, we explore the use of local vision-language models (VLMs), combined with large language models (LLMs), and compare them to a supervised TimeSFormer baseline. Using a simulated dataset comprising 13.26 hours of newborn resuscitation videos, we evaluate several zero-shot VLM-based strategies and fine-tuned VLMs with classification heads, including Low-Rank Adaptation (LoRA). Our results suggest that small (local) VLMs struggle with hallucinations, but when fine-tuned with LoRA, the results reach F1 score at 0.91, surpassing the TimeSformer results of 0.70.

视觉语言模型动作识别生成式AI医疗视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。