Lance统一建模图文视频,用协同训练实现理解生成兼优。
Lance: Unified Multimodal Modeling by Multi-Task Synergy

- 双流专家混合架构,共享序列联合学习上下文
- 图像视频生成能力超越现有开源模型,理解力强
- 适合需要统一处理多模态内容的开发者和研究者
我们提出Lance,一种轻量级原生统一模型,支持图像与视频的多模态理解、生成与编辑。不同于依赖模型规模扩展或以文本-图像为主的设计,Lance通过协作式多任务训练探索实用的统一多模态建模范式。其基于两大核心原则:统一上下文建模与解耦能力路径。Lance从头训练,采用双流专家混合架构处理共享的交错多模态序列,实现联合上下文学习的同时解耦理解与生成路径。我们进一步引入模态感知的旋转位置编码,缓解异质视觉标记间的干扰并增强跨任务对齐。训练阶段采用分阶段多任务策略,结合能力导向目标与自适应数据调度,强化语义理解与视觉生成性能。实验表明,Lance在图像与视频生成上显著优于现有开源统一模型,同时保持强大的多模态理解能力。
原文摘要 · Abstract (English)
We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, Lance explores a practical paradigm for unified multimodal modeling via collaborative multi-task training. It is grounded in two core principles: unified context modeling and decoupled capability pathways. Specifically, Lance is trained from scratch and employs a dual-stream mixture-of-experts architecture on shared interleaved multimodal sequences, enabling joint context learning while decoupling the pathways for understanding and generation. We further introduce modality-aware rotary positional encoding to mitigate interference among heterogeneous visual tokens and boost cross-task alignment. During training, Lance adopts a staged multi-task training paradigm with capability-oriented objectives and adaptive data scheduling to strengthen both semantic comprehension and visual generation performance. Experimental results demonstrate that Lance substantially outperforms existing open-source unified models in image and video generation, while retaining strong multimodal understanding capabilities. The homepage is available at https://lance-project.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。