arXiv:2501.04150cs.CV2025-01被引 6

对比大模型与小模型在多模态任务中的表现边界。

Benchmarking Large and Small MLLMs

  • 构建系统性评测框架,覆盖通用能力与工业、汽车等真实场景。
  • 小模型在特定任务上接近大模型表现,复杂推理任务仍显著落后。
  • 揭示大小模型共有的失败模式,指明未来改进方向。

大型多模态语言模型(MLLMs)如GPT-4V和GPT-4o在理解与生成多模态内容方面取得显著进展,展现出优异的跨任务性能。然而,其部署面临推理慢、计算成本高、难以在设备端运行等问题。相比之下,小型MLLMs(如LLava系列和Phi-3-Vision)凭借更快的推理速度、更低的部署成本,以及处理领域特定任务的能力,成为有前景的替代方案。尽管如此,大模型与小模型的能力边界尚不清晰。本文通过系统化、全面的评估,对大小型MLLMs进行对比测试,涵盖物体识别、时间推理、多模态理解等通用能力,以及工业和汽车等实际应用。结果表明,小模型在特定场景下可达到与大模型相当的性能,但在需要深度推理或细微理解的复杂任务中仍存在明显差距。此外,我们识别出大小模型共有的失败案例,揭示了当前先进模型仍难以应对的领域。期望本研究能推动学术界进一步突破多模态模型的性能边界,提升其在多样化应用场景中的实用性与有效性。

原文摘要 · Abstract (English)

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their deployment faces significant challenges, including slow inference, high computational cost, and impracticality for on-device applications. In contrast, the emergence of small MLLMs, exemplified by the LLava-series models and Phi-3-Vision, offers promising alternatives with faster inference, reduced deployment costs, and the ability to handle domain-specific scenarios. Despite their growing presence, the capability boundaries between large and small MLLMs remain underexplored. In this work, we conduct a systematic and comprehensive evaluation to benchmark both small and large MLLMs, spanning general capabilities such as object recognition, temporal reasoning, and multimodal comprehension, as well as real-world applications in domains like industry and automotive. Our evaluation reveals that small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks requiring deeper reasoning or nuanced understanding. Furthermore, we identify common failure cases in both small and large MLLMs, highlighting domains where even state-of-the-art models struggle. We hope our findings will guide the research community in pushing the quality boundaries of MLLMs, advancing their usability and effectiveness across diverse applications.

多模态模型大模型小模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。