InteractiveOmni让模型同时听懂语音、看懂视频,还能长对话记忆。
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- 统一融合视觉、音频、语言和语音生成模块,支持多模态交互。
- 4B版本性能接近7B大模型,8B版仅用一半参数就保留97%能力。
- 专为长对话设计,适合智能助手、人机交互等实际应用。
我们提出InteractiveOmni,一个参数规模为4B至8B的开源统一多模态大模型,用于音视频多轮交互。该模型整合视觉编码器、音频编码器、语言模型与语音解码器,实现理解与生成一体化。通过分阶段训练策略,包括多模态预训练、语音对话后训练及视听交互优化,提升跨模态能力。为增强长期对话能力,我们精心构建多轮对话数据集,并建立多模态多轮记忆基准与多轮语音交互基准进行评估。实验表明,InteractiveOmni显著优于现有开源模型,在图像、音频、视频理解及语音生成任务上均达同类规模领先水平。其中,InteractiveOmni-4B在通用评测中表现媲美更大模型Qwen2.5-Omni-7B,而InteractiveOmni-8B仅需50%参数量即可保持97%性能,是下一代智能交互系统的理想基础。
原文摘要 · Abstract (English)
We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive omni-modal understanding and speech generation capabilities. To achieve this, we integrate the vision encoder, audio encoder, large language model, and speech decoder into a unified model for understanding and generation tasks. We design a multi-stage training strategy to ensure robust cross-modal capabilities, including pre-training for omni-modal understanding, followed by post-training with speech conversation and audio-visual interaction. To enable human-like long-term conversational ability, we meticulously curate a multi-turn training dataset that enhances the model's ability to handle complex and multi-turn interactions. To effectively evaluate the multi-turn memory and speech interaction capabilities, we construct the multi-modal multi-turn memory benchmark and the multi-turn speech interaction benchmark. Experiments demonstrate that InteractiveOmni significantly outperforms leading open-source models and provides a more intelligent multi-turn audio-visual experience, particularly in its long-term memory capabilities. Notably, InteractiveOmni-4B is comparable to the much larger model like Qwen2.5-Omni-7B on general benchmarks, and it can retain 97% of the performance of the InteractiveOmni-8B while utilizing only 50% of the model size. Achieving state-of-the-art results against similarly sized models across image, audio, video understanding, and speech generation tasks, InteractiveOmni is an accessible, open-source foundation for next-generation intelligent interactive systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。