自动挑选最适合任务的视觉语言模型,省时省力。
Mordal: Automated Pretrained Model Selection for Vision Language Models
- 基于高效搜索策略,自动筛选最优视觉语言模型。
- 相比传统方法,节省8.9至11.6倍的GPU计算时间。
- 在多类任务上表现优于现有最佳选择方法。
将多模态信息融入大语言模型是提升其对非文本数据理解能力的有效途径,使模型能够完成多模态任务。视觉语言模型(VLMs)因其在医疗、机器人和无障碍等领域的广泛应用,成为增长最快的多模态模型类别。然而,尽管现有VLM在不同基准测试中表现出色,但它们均由人工专家手工设计,缺乏自动化框架实现任务特定模型构建。本文提出Mordal,一个全自动的多模态模型搜索框架,可在无需人工干预的情况下高效找到适用于用户定义任务的最佳VLM。该框架通过减少候选模型数量并缩短每个候选的评估时间来实现高效搜索。实验表明,与网格搜索相比,Mordal在完成相同任务时可节省8.9×–11.6×的GPU小时数;在多种任务上,其加权肯德尔等级相关系数(Kendall's τ)平均高出69%以上,显著优于当前最先进的模型选择方法。
原文摘要 · Abstract (English)
Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing category of multimodal models because of their many practical use cases, including in healthcare, robotics, and accessibility. Unfortunately, even though different VLMs in the literature demonstrate impressive visual capabilities in different benchmarks, they are handcrafted by human experts; there is no automated framework to create task-specific multimodal models. We introduce Mordal, an automated multimodal model search framework that efficiently finds the best VLM for a user-defined task without manual intervention. Mordal achieves this both by reducing the number of candidates to consider during the search process and by minimizing the time required to evaluate each remaining candidate. Our evaluation shows that Mordal can find the best VLM for a given problem using $8.9\times$--$11.6\times$ lower GPU hours than grid search. We have also discovered that Mordal achieves about 69\% higher weighted Kendall's $τ$ on average than the state-of-the-art model selection method across diverse tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。