arXiv:2504.14391cs.CV2025-04被引 7

用YouTube医学视频训练模型,效果出人意料好

How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?

  • 用人工筛选的1031小时医学视频+问答对微调视觉语言模型
  • 2B模型在视频任务上提升98.7%,7B模型在手术视频上提升52.1%
  • 适合医学多模态研究者和医疗AI开发者

公开的生物医学视频(如YouTube内容)是医学生的重要教育资源,其特点为混合影像、解说与上下文框架,非标准化但教学性强。本文构建了包含1031小时视频-字幕与问答对的OpenBiomedVid数据集,通过多阶段人机协同流程筛选。实验表明,尽管视频内容非标准化,微调后的Qwen-2-VL模型仍显著提升性能:2B模型在视频任务上提升98.7%,图像任务提升71.2%;7B模型在视频任务上提升37.09%,图像任务提升11.2%,文本任务略有下降(-2.7%)。为弥补评估数据缺失,本文还引入两个专家标注的新基准MIMICEchoQA和SurgeryVideoQA。在这些更标准的数据集上,2B模型分别提升99.1%和98.1%,7B模型分别提升22.5%和52.1%,证明模型具备良好泛化能力。

原文摘要 · Abstract (English)

Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVi, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7% on video tasks, 71.2% on image tasks, and 0.2% on text tasks. The 7B model shows improvements of 37.09% on video and 11.2% on image tasks, with a slight degradation of 2.7% on text tasks compared to their respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1% and 98.1%, while the 7B model shows gains of 22.5% and 52.1%, respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs.

医学AI视觉语言模型视频理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。