arXiv:2602.15257cs.CVcs.AI2026-02被引 3

首次系统研究34万上下文长文档视觉模型训练方法,提升长文本问答性能。

How to Train Your Long-Context Visual Document Model

  • 按评估长度训练比用更长上下文效果更好
  • 引入页码索引显著提升长文档理解能力
  • 验证了视觉长上下文训练可反向提升文本任务表现

本文首次系统开展针对长达344K上下文的长文档视觉语言模型大规模训练研究,聚焦长文档视觉问答任务,并评估其在长文本上的迁移性能。尽管已有多个开源大模型(如Qwen3 VL、GLM 4.5/6V)具备强能力,但其训练方案与数据流程不可复现。本研究对24B和32B参数模型进行持续预训练、监督微调与偏好优化的系统实验,结合大量长上下文评估与消融分析,实现MMLongBenchDoc基准上双规模模型的当前最优表现。关键发现包括:(i) 按评估上下文长度训练优于使用更长上下文;(ii) 在训练与评估中加入页码索引带来显著性能提升;(iii) 自研合成数据管道支持通过持续预训练与微调实现模型自优化;(iv) 首次验证视觉长上下文训练可正向迁移至长文本任务。此外,发布经人工修正的MMLBD-C版本,减少原基准中的错误与低质量样本。

原文摘要 · Abstract (English)

We present the first comprehensive, large-scale study of training long-context vision language models up to 344K context, targeting long-document visual question answering with measured transfer to long-context text. While several such strong are open-weight, namely Qwen3 VL and GLM 4.5/6V, their training recipes and data pipelines are not reproducible. We systematically study continued pretraining, supervised finetuning, and preference optimization for 24B and 32B parameter models, backed by extensive LC evaluations and ablations to bridge this gap, and achieve state-of-the-art performance on MMLongBenchDoc for both parameter scales. In addition to this, our key findings include: (i) training on context lengths that match evaluation context lengths outperforms training on longer contexts, (ii) training and evaluating with page indices provides a simple, high-impact boost to long-document performance, (iii) our synthetic data pipelines enable self-improvement via continued pretraining and supervised finetuning, and (iv) we extend the known text-to-visual long context transfer to the reverse, showing that visual long context training transfers to long-context text performance. We also release MMLBD-C, a manually corrected version of MMLongBenchDoc to reduce erroneous and low quality examples in the benchmark.

长上下文视觉问答模型训练多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。