Qwen3.5-Omni支持超长上下文,可直接根据音视频指令编程。
Qwen3.5-Omni Technical Report

- 采用混合注意力MoE架构,支持超长序列高效推理。
- 在215项任务中达顶尖水平,音频理解优于Gemini-3.1 Pro。
- 首次实现音视频指令直接生成代码,适合多模态交互研究者。
本文介绍Qwen3.5-Omni,Qwen-Omni系列最新进展。该模型参数量达数百亿级,支持256k上下文长度。基于包含异构图文对及超1亿小时音视频内容的海量数据训练,展现出强大的多模态能力。Qwen3.5-Omni-plus在215个音频与音视频理解、推理及交互子任务中表现领先,关键音频任务超越Gemini-3.1 Pro,综合音视频理解持平。其采用混合注意力专家(Hybrid Attention MoE)架构,分别应用于Thinker与Talker模块,实现长序列高效推理。支持超过10小时音频理解与400秒720P视频(1 FPS)处理。针对流式语音合成中因文本与语音分词器效率差异导致的不自然问题,提出ARIA动态对齐机制,显著提升语音稳定性与韵律,延迟影响极小。模型扩展至10种语言,具备类人情感表达能力。在音视频定位方面,能生成精准时间同步的结构化字幕与自动场景分割。尤为关键的是,观察到新型能力:直接根据音视频指令进行编码,称为音频-视觉氛围编码(Audio-Visual Vibe Coding)。
原文摘要 · Abstract (English)
In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-plus achieves SOTA results across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding. Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (MoE) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis, often caused by encoding efficiency discrepancies between text and speech tokenizers, we introduce ARIA. ARIA dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding and speech generation across 10 languages with human-like emotional nuance. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。