arXiv:2509.18189cs.CVcs.AI2025-09被引 4

百度推出多模态大模型Qianfan-VL,强化领域能力且保持通用性能。

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

  • 采用分阶段渐进训练与高精度数据合成,提升领域适应性。
  • 在OCR和文档理解上表现突出,如DocVQA达94.75%准确率。
  • 适合企业级应用,支持长链推理与大规模部署。

我们提出Qianfan-VL,一系列参数规模从30亿到700亿的多模态大语言模型,通过创新的领域增强技术实现领先性能。该方法结合分阶段渐进训练与高精度数据合成流程,在保持强大通用能力的同时显著提升领域特定能力。Qianfan-VL在通用基准上表现媲美顶尖开源模型,在CCBench、SEEDBench IMG、ScienceQA和MMStar等任务中达到最优水平。其领域增强策略在OCR与文档理解方面优势明显,公开基准测试中OCRBench得分873,DocVQA达94.75%,内部评估亦验证效果。其中,Qianfan-VL-8B与70B版本具备长链思维能力,在数学推理(MathVista 78.6%)与逻辑推理任务中表现优异。所有模型均基于百度昆仑P800芯片完成训练,单任务在5000张芯片上实现超过90%的缩放效率。本工作为多模态领域增强模型的开发提供了可落地的方法论,适用于多样化企业部署场景。

原文摘要 · Abstract (English)

We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement techniques. Our approach employs multi-stage progressive training and high-precision data synthesis pipelines, which prove to be critical technologies for enhancing domain-specific capabilities while maintaining strong general performance. Qianfan-VL achieves comparable results to leading open-source models on general benchmarks, with state-of-the-art performance on benchmarks such as CCBench, SEEDBench IMG, ScienceQA, and MMStar. The domain enhancement strategy delivers significant advantages in OCR and document understanding, validated on both public benchmarks (OCRBench 873, DocVQA 94.75%) and in-house evaluations. Notably, Qianfan-VL-8B and 70B variants incorporate long chain-of-thought capabilities, demonstrating superior performance on mathematical reasoning (MathVista 78.6%) and logical inference tasks. All models are trained entirely on Baidu's Kunlun P800 chips, validating the capability of large-scale AI infrastructure to train SOTA-level multimodal models with over 90% scaling efficiency on 5000 chips for a single task. This work establishes an effective methodology for developing domain-enhanced multimodal models suitable for diverse enterprise deployment scenarios.

多模态大模型文档理解OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。