arXiv:2412.08443cs.CVcs.MM2024-12被引 11

POINTS1.5提升多分辨率图像处理与中文理解能力,适合真实场景应用。

POINTS1.5: Building a Vision-Language Model towards Real World Applications

  • 采用动态高分辨率视觉编码器,支持任意尺寸图像输入。
  • 通过自建中英文数据集,显著增强中文任务表现,7B模型仅用40亿词训练。
  • 设计严格筛选方法优化指令微调数据,性能在同类模型中领先。

近年来,视觉语言模型在多项任务上取得显著进展,如光学字符识别和复杂图表分析。在此基础上,我们提出新模型 POINTS1.5,旨在更好支持各类真实世界应用。POINTS1.5 是 POINTS1.0 的升级版,包含三项关键改进:(i)将原始固定分辨率的 CLIP 视觉编码器替换为支持原生动态高分辨率的 NaViT 风格编码器,使模型可直接处理任意分辨率图像,无需分块;(ii)引入双语支持,大幅提升中文理解能力。由于开源中文视觉语言数据稀缺,我们从互联网收集大量图像,并结合人工与自动方式完成标注;(iii)提出一套严格的视觉指令微调数据集过滤方法,全面评估后选择最优策略构建最终数据集。得益于这些创新,POINTS1.5 显著超越 POINTS1.0,展现出强大实用性。特别地,POINTS1.5-7B 模型仅使用不到 40 亿个训练词元,在参数少于 100 亿的模型中,位列 OpenCompass 领先榜首位。

原文摘要 · Abstract (English)

Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character recognition and complex diagram analysis. Building on this trend, we introduce a new vision-language model, POINTS1.5, designed to excel in various real-world applications. POINTS1.5 is an enhancement of POINTS1.0 and incorporates several key innovations: i) We replace the original CLIP vision encoder, which had a fixed image resolution, with a NaViT-style vision encoder that supports native dynamic high resolution. This allows POINTS1.5 to process images of any resolution without needing to split them into tiles. ii) We add bilingual support to POINTS1.5, significantly enhancing its capability in Chinese. Due to the scarcity of open-source Chinese datasets for vision-language models, we collect numerous images from the Internet and annotate them using a combination of manual and automatic methods. iii) We propose a set of rigorous filtering methods for visual instruction tuning datasets. We comprehensively evaluate all these filtering methods, and choose the most effective ones to obtain the final visual instruction tuning set. Thanks to these innovations, POINTS1.5 significantly outperforms POINTS1.0 and demonstrates strong performance across a range of real-world applications. Notably, POINTS1.5-7B is trained on fewer than 4 billion tokens and ranks first on the OpenCompass leaderboard among models with fewer than 10 billion parameters

视觉语言模型中文理解多分辨率真实应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。