开源韩英双语视觉语言模型,支持图文理解与生成。
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
- 分步训练保留主干知识,同步学习语言与视觉信息。
- 在双语图文任务中表现优于同规模模型。
- 支持定位、指代与OCR,适合真实场景应用。
本文介绍一个开源的韩英双语视觉语言模型(VLM)VARCO-VISION。通过分步训练策略,模型在学习语言与视觉信息的同时,有效保留了主干模型的知识。相比同规模模型,VARCO-VISION在需双语图文理解与生成的多种任务中表现优异。该模型还具备定位、指代和光学字符识别(OCR)能力,拓展了其在实际场景中的应用潜力。除模型外,我们还发布了五个韩语评估数据集,包括四个闭集与一个开集基准。我们期望这一里程碑能为致力于训练视觉语言模型的研究者提供更多机会。VARCO-VISION 已在 Hugging Face 上开源:https://huggingface.co/NCSOFT/VARCO-VISION-14B。
原文摘要 · Abstract (English)
In this paper, we introduce an open-source Korean-English vision-language model (VLM), VARCO-VISION. We incorporate a step-by-step training strategy that allows a model learn both linguistic and visual information while preserving the backbone model's knowledge. Our model demonstrates outstanding performance in diverse settings requiring bilingual image-text understanding and generation abilities compared to models of similar size. VARCO-VISION is also capable of grounding, referring, and OCR, expanding its usage and potential applications for real-world scenarios. In addition to the model, we release five Korean evaluation datasets, including four closed-set and one openset benchmarks. We anticipate that our milestone will broaden the opportunities for AI researchers aiming to train VLMs. VARCO-VISION is available at https://huggingface.co/NCSOFT/VARCO-VISION-14B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。