arXiv:2606.01049cs.CL2026-06

用原文图注重建医学多模态数据,提升模型医疗理解能力。

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

论文配图:Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining
图 1 · 摘自论文原文
  • 通过图注定位重构图文序列,确保图文关联真实可追溯。
  • 构建9.63B词元医学多模态语料,使模型在医疗任务上提升1.46分。
  • 适合做医学大模型预训练或需精准图文对齐的研究者使用。

医学图表的理解不仅依赖图注,还需正文中的讨论支持。当前多模态数据常将图表简化为孤立的图像-文本对,丢失关键上下文。现有方法要么忽略上下文,要么随意附加,导致图文关联不实、论述断裂。本文提出上下文锚定重建框架,将PubMed Central开放获取(PMC-OA)文献转化为具有参照一致性的交错序列:恢复图注与原文,仅通过文章内原生图注连接上下文,修复非连续内容,剔除无支持的图像。基于此,构建了9.63B词元的医学多模态持续预训练语料(PMC-InterCPT),经文本质量筛选与证据感知分配后,用于生成式医学多模态大模型(MLLM)的持续预训练。在固定监督微调下,相比等量词元的原始数据控制组,其使Qwen3.5-4B-Base在医疗平均得分上提升1.46分,通用/科学平均得分提升3.11分,并超越42%更大的原始数据训练结果。效果扩展至其他模型。受控消融实验表明,上下文锚定重建是核心,而非简单拼接或单纯扩大数据量。

原文摘要 · Abstract (English)

Biomedical figures are explained not by captions alone but by body-text passages that discuss them. Yet current multimodal corpora typically reduce figures to isolated image-caption pairs, discarding this crucial context. Existing pipelines either omit this context or append it without enforcing the figure references that support each attachment, which can create unsupported image-text attachments and incoherent discourse. We introduce context-grounded reconstruction, a source-grounded framework that converts PubMed Central Open Access (PMC-OA) records into referentially coherent interleaved sequences. It recovers captions and source text, attaches context only through article-native figure references, repairs non-contiguous context, and prunes unsupported images. Starting from these reconstructed sequences, PMC-InterCPT first filters records for text quality and medical relevance, then applies evidence-aware allocation to form a 9.63B-token corpus for continued pretraining (CPT) of generative medical MLLMs. With fixed supervised fine-tuning (SFT), PMC-InterCPT improves Qwen3.5-4B-Base by 1.46 medical-average points and 3.11 general/scientific-average points over a token-matched raw source control, and surpasses a 42% larger raw-data run. Gains transfer to Qwen3.5-2B-Base and LLaVA-OneVision-1.5-4B-Base. Controlled ablations show that context-grounded reconstruction, rather than simply appending article context or scaling raw data, is central to useful biomedical multimodal CPT.

医学多模态持续预训练图文对齐数据重构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。