arXiv:2511.21631cs.CVcs.AI2025-11被引 2.1k

Qwen3-VL支持256K长上下文,多模态理解更强。

Qwen3-VL Technical Report

  • 采用交错式建模与深层视觉对齐,提升图文视频理解能力
  • 在256K上下文下保持精准跨模态检索与推理,优于同类模型
  • 适合需要长文档分析、视频理解的智能系统研发

我们介绍 Qwen3-VL,这是目前 Qwen 系列中最强大的多模态模型,在广泛的多模态基准测试中表现优异。它原生支持长达 256K token 的交错文本、图像和视频上下文,包含稠密型(2B/4B/8B/32B)和混合专家型(30B-A3B/235B-A22B)变体,以适应不同延迟-质量权衡需求。Qwen3-VL 实现三大核心能力:(i) 显著增强的纯文本理解,部分场景超越同级别文本基座模型;(ii) 强大的长上下文理解能力,原生支持 256K token 的文本及交错多模态输入,实现长文档与视频中信息的忠实保留、检索与交叉引用;(iii) 在单图、多图及视频任务上具备先进多模态推理能力,在 MMMU 及视觉数学基准(如 MathVista、MathVision)中表现领先。架构上引入三项关键升级:(i) 改进的交错式 MRoPE,强化图像与视频中的时空建模;(ii) DeepStack 集成,有效利用多层级 ViT 特征,增强视觉-语言对齐;(iii) 基于文本的时间对齐机制,从 T-RoPE 进化为显式文本时间戳对齐,实现更精确的时间定位。在同等令牌预算与延迟约束下,无论稠密或混合专家(MoE)架构,均实现更优性能。我们预期 Qwen3-VL 将成为图像引导推理、智能体决策与多模态代码智能等真实工作流的基础引擎。

原文摘要 · Abstract (English)

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.

多模态长上下文视觉语言模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。