arXiv:2512.02566cs.CVcs.AI2025-12

将医学文献中的多面板图按层级拆解,提升视觉语言模型的细粒度理解能力。

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

  • 从文献图中提取多层级结构,生成图、面板、局部区域三级对齐数据。
  • 仅用少量文献图即获得更优监督信号,显著提升模型性能。
  • 适合医学图像分析、病理诊断等需精细定位的场景。

医学视觉语言模型的研究日益增多,现有方法多依赖网络规模的科学数据,但通常将丰富的科学图表与文本压缩为粗粒度的图级配对,丢失了临床医生在放大局部结构时依赖的细粒度对应关系。为此,我们提出Panel2Patch,一种新型数据处理流程,从现有生物医学科学文献中挖掘多层次结构——即多面板、标记密集型图表及其周边文本,并转化为多粒度监督信号。给定科学图表与标题,Panel2Patch解析布局、面板及视觉标记,构建图、面板和局部区域三个层级的对齐视觉-语言对,保留局部语义而非将每张图视为单一样本。基于此层次化语料库,我们设计了一种粒度感知的预训练策略,统一来自粗粒度教学描述到细粒度区域聚焦短语的异构目标。仅使用少量文献图表,Panel2Patch提取的有效监督信号远超以往流程,使模型在更少预训练数据下实现显著更好的性能。

原文摘要 · Abstract (English)

There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. However, current biomedical vision-language pretraining typically compresses rich scientific figures and text into coarse figure-level pairs, discarding the fine-grained correspondences that clinicians actually rely on when zooming into local structures. To tackle this issue, we introduce Panel2Patch, a novel data pipeline that mines hierarchical structure from existing biomedical scientific literature, i.e., multi-panel, marker-heavy figures and their surrounding text, and converts them into multi-granular supervision. Given scientific figures and captions, Panel2Patch parses layouts, panels, and visual markers, then constructs hierarchical aligned vision-language pairs at the figure, panel, and patch levels, preserving local semantics instead of treating each figure as a single data sample. Built on this hierarchical corpus, we develop a granularity-aware pretraining strategy that unifies heterogeneous objectives from coarse didactic descriptions to fine region-focused phrases. By applying Panel2Patch to only a small set of the literature figures, we extract far more effective supervision than prior pipelines, enabling substantially better performance with less pretraining data.

视觉语言医学图像细粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。