Omni模型通过跨模态推理提升多模态理解与生成能力。
Context Unrolling in Omni Models

- 在多种模态上联合训练,实现跨模态上下文展开推理。
- 在多模态生成与理解任务上表现优异,支持文本、图像、视频和3D建模生成。
- 适合需要复杂多模态推理的场景,如跨模态内容创作。
我们提出Omni,一个原生在多种模态(包括文本、图像、视频、3D几何和隐式表示)上训练的统一多模态模型。实验发现,这种训练方式催生了上下文展开(Context Unrolling)机制,使模型能在输出预测前显式地跨多个模态表示进行推理。该过程有效聚合异构模态间的互补信息,更准确逼近共享的多模态知识流形,从而提升下游任务的推理保真度。Omni在多模态生成与理解基准测试中均表现强劲,展现出先进的多模态推理能力,包括在上下文中生成文本、图像、视频及3D几何体的能力。
原文摘要 · Abstract (English)
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。