用智能体协作让模型跨模态推理,无需重训即可理解图文音视频
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
- 通过主控智能体协调不同模态专用模型完成任务
- 在多模态基准上达到当前最佳性能,尤其擅长复杂跨模态推理
- 模块化设计易扩展,适合需要灵活多模态理解的场景
多模态大语言模型虽具强大能力,但仍受限于固定模态组合,且需大量对齐数据进行昂贵微调。构建可处理文本、图像、音频、视频的全模态模型仍不现实,且缺乏稳健的推理支持。本文提出Agent-Omni框架,通过主控智能体协调现有基础模型,实现无需重训的灵活多模态推理。主智能体解析用户意图,将子任务分发给特定模态智能体,并整合其输出生成连贯回答。在文本、图像、音频、视频及全模态基准上的大量实验表明,Agent-Omni持续取得领先表现,尤其在需要复杂跨模态推理的任务中。其智能体架构可无缝集成专业基础模型,确保对多样化输入的适应性,同时保持透明与可解释性。此外,该框架具有模块化特性,易于扩展,未来可随更强模型的出现而升级。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text, images, audio, and video remains impractical and lacks robust reasoning support. In this paper, we propose an Agent-Omni framework that coordinates existing foundation models through a master-agent system, enabling flexible multimodal reasoning without retraining. The master agent interprets user intent, delegates subtasks to modality-specific agents, and integrates their outputs into coherent responses. Extensive experiments across text, image, audio, video, and omni benchmarks show that Agent-Omni consistently achieves state-of-the-art performance, particularly on tasks requiring complex cross-modal reasoning. Its agent-based design enables seamless integration of specialized foundation models, ensuring adaptability to diverse inputs while maintaining transparency and interpretability. In addition, the framework is modular and easily extensible, allowing future improvements as stronger models become available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。