arXiv:2501.04184cs.CV2025-01NeurIPS被引 4

用医学视频字幕和鼠标轨迹构建大规模医疗图文数据集

MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

  • 从医学教学视频提取图文对与像素级鼠标轨迹
  • 100万图像带空间标注,11.8万视频含时序对齐文本
  • 新模型在12个领域均超越现有最佳水平

多模态模型依赖大量数据。虽然自然图像数据丰富,但医疗图像数据难以获得同等规模。为实现医疗图像的规模化表征学习,我们转向YouTube平台,该平台拥有大量开源医学教学视频。我们构建了MedicalNarratives数据集,包含470万组医疗图像-文本对,其中100万样本带有密集的空间标注(包括像素级鼠标轨迹和边界框),以及11.8万段以轨迹事件为中心、文本对齐的视频,支持跨帧时空定位。类似于“自言自语”研究中教师边操作鼠标边讲解,该数据集中100万张图像包含像素级的局部鼠标轨迹,建立了文本与图像像素间的时空关联。为评估其价值,我们基于该数据集,采用类CLIP目标训练了GenMedClip模型,覆盖12个医学领域。在新构建的医疗影像基准测试中,GenMedClip在所有12个领域均优于先前最优模型。

原文摘要 · Abstract (English)

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a large reservoir of open-source medical pedagogical videos. We curate MedicalNarratives, a dataset 4.7M medical image-text pairs, with 1M samples containing dense annotations in the form of spatial traces (and bounding boxes), and 118K videos centered on the trace event (with aligned text), enabling spatiotemporal grounding beyond single frames. Similar to $\textit{think-aloud}$ studies where instructors speak while hovering their mouse cursor movements over relevant image regions, 1M images in MedicalNarratives contains localized mouse traces in image pixels, creating a spatial and temporal association between the text and pixels. To evaluate the utility of MedicalNarratives, we train GenMedClip with a CLIP-like objective using our dataset spanning 12 medical domains. GenMedClip outperforms previous state-of-the-art models on all 12 domains on a newly constructed medical imaging benchmark. $\href{https://huggingface.co/datasets/wisdomik/MedicalNarratives}{[Data]}$

医疗视觉多模态数据集时空对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。