构建开源多模态大模型,提升视觉与音频理解能力
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- 设计多模态对齐网络与时间编码机制,增强跨模态表征
- 用2400万条对话数据训练,性能超越同类模型19%以上
- 适合机器人、医疗AI等需要多模态感知的场景
推进机器智能需具备多模态感知能力。我们提出OmniVinci,一个开源的强多模态大模型。通过精心设计模型架构与数据构建流程,提出三项创新:(i) OmniAlignNet,在共享多模态潜在空间中强化视觉与音频嵌入对齐;(ii) 时序嵌入分组,捕捉视觉与音频信号间的相对时间对齐;(iii) 约束旋转时间嵌入,编码多模态表示中的绝对时间信息。构建数据编排与合成流水线,生成2400万条单模态与多模态对话。发现模态在感知与推理中相互增强。OmniVinci在DailyOmni(跨模态理解)上比Qwen2.5-Omni高出+19.05,在MMAR(音频)上+1.7,在Video-MME(视觉)上+3.9,仅使用0.2万亿训练标记,为Qwen2.5-Omni的1.2万亿的1/6。最终在机器人、医疗AI、智能工厂等下游任务中验证多模态优势。
原文摘要 · Abstract (English)
Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curation. For model architecture, we present three key innovations: (i) OmniAlignNet for strengthening alignment between vision and audio embeddings in a shared omni-modal latent space; (ii) Temporal Embedding Grouping for capturing relative temporal alignment between vision and audio signals; and (iii) Constrained Rotary Time Embedding for encoding absolute temporal information in omni-modal embeddings. We introduce a curation and synthesis pipeline that generates 24M single-modal and omni-modal conversations. We find that modalities reinforce one another in both perception and reasoning. Our model, OmniVinci, outperforms Qwen2.5-Omni with +19.05 on DailyOmni (cross-modal understanding), +1.7 on MMAR (audio), and +3.9 on Video-MME (vision), while using just 0.2T training tokens - a 6 times reduction compared to Qwen2.5-Omni's 1.2T. We finally demonstrate omni-modal advantages in downstream applications spanning robotics, medical AI, and smart factory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。