arXiv:2510.14979cs.CVcs.AI2025-10中稿 · ICLR被引 14

提出新型统一视觉语言模型NEO,实现像素与文本的深度融合。

From Pixels to Words -- Towards Native Vision-Language Primitives at Scale

  • 从头构建统一模型,融合视觉与语言模块优势
  • 基于390万图文对训练,显著降低模态冲突
  • 开源可复用组件,推动低成本高效研究

原生视觉语言模型(VLM)正逐渐挑战传统模块化架构,但其发展仍受两大瓶颈制约:一是原生模型与模块化模型的根本差异及突破边界;二是如何让研究更开放、易用以加速领域进步。本文厘清这些挑战,提出构建原生VLM的核心原则:(i)在共享语义空间中有效对齐图像像素与文字表示;(ii)无缝整合视觉与语言模块的优势;(iii)内嵌多种跨模态特性,支持统一编码、对齐与推理。据此,我们推出全新原生VLM家族NEO,基于390万图文样本从零训练,通过密集统一结构缓解内部模态冲突,在多种真实场景中逼近顶级模块化模型性能。NEO作为可扩展、高性能原生模型的基石,配套丰富可复用组件,构建低成本、可拓展的研究生态。代码与模型已公开于 https://github.com/EvolvingLMMs-Lab/NEO。

原文摘要 · Abstract (English)

The edifice of native Vision-Language Models (VLMs) has emerged as a rising contender to typical modular VLMs, shaped by evolving model architectures and training paradigms. Yet, two lingering clouds cast shadows over its widespread exploration and promotion: (-) What fundamental constraints set native VLMs apart from modular ones, and to what extent can these barriers be overcome? (-) How to make research in native VLMs more accessible and democratized, thereby accelerating progress in the field. In this paper, we clarify these challenges and outline guiding principles for constructing native VLMs. Specifically, one native VLM primitive should: (i) effectively align pixel and word representations within a shared semantic space; (ii) seamlessly integrate the strengths of formerly separate vision and language modules; (iii) inherently embody various cross-modal properties that support unified vision-language encoding, aligning, and reasoning. Hence, we launch NEO, a novel family of native VLMs built from first principles, greatly narrowing the gap with top-tier modular counterparts across diverse real-world scenarios. With 390M image-text examples, NEO efficiently develops visual perception from scratch while mitigating vision-language conflicts inside a dense and monolithic model crafted from our elaborate primitives. We position NEO as a cornerstone for scalable and powerful native VLM development, paired with a rich set of reusable components that foster a cost-effective and extensible ecosystem. Our code and models are publicly available at: https://github.com/EvolvingLMMs-Lab/NEO.

视觉语言模型统一架构开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。