arXiv:2509.10105cs.CVcs.CL2025-09被引 5

双语视觉语言模型支持图文空间定位,适配韩英双语文档理解。

VARCO-VISION-2.0 Technical Report

  • 四阶段训练+高效内存技术,提升多模态对齐与语言能力。
  • 14B模型在OpenCompass榜单中位列同类模型第8,1.7B版可本地部署。
  • 擅长处理图表、表格等复杂输入,输出文字内容及位置信息。

我们推出VARCO-VISION-2.0,一个开源权重的韩英双语视觉语言模型,相较于前代14B版本能力更优。该模型支持多图像理解,适用于文档、图表和表格等复杂输入,并通过预测文本内容及其空间位置实现布局感知的OCR。采用四阶段课程训练结合内存高效技术,增强了多模态对齐,同时保持核心语言能力并借助偏好优化提升安全性。大量基准测试显示,其空间定位能力强,双语表现具有竞争力;14B模型在与同规模模型对比的OpenCompass VLM排行榜中位列第8。此外,还发布了适合设备端部署的1.7B轻量版本。两个变体均已在Hugging Face上线。

原文摘要 · Abstract (English)

We introduce VARCO-VISION-2.0, an open-weight bilingual vision-language model (VLM) for Korean and English with improved capabilities compared to the previous model VARCO-VISION-14B. The model supports multi-image understanding for complex inputs such as documents, charts, and tables, and delivers layoutaware OCR by predicting both textual content and its spatial location. Trained with a four-stage curriculum with memory-efficient techniques, the model achieves enhanced multimodal alignment, while preserving core language abilities and improving safety via preference optimization. Extensive benchmark evaluations demonstrate strong spatial grounding and competitive results for both languages, with the 14B model achieving 8th place on the OpenCompass VLM leaderboard among models of comparable scale. Alongside the 14B-scale model, we release a 1.7B version optimized for on-device deployment. We believe these models advance the development of bilingual VLMs and their practical applications. Two variants of VARCO-VISION-2.0 are available at Hugging Face: a full-scale 14B model and a lightweight 1.7B model.

双语VLM布局感知OCR多模态理解轻量化部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。