在统一实验中评估模型合并效果,揭示各方法优劣与实际需求
Realistic Evaluation of Model Merging for Compositional Generalization
- 统一实验环境对比多种合并方法的性能与成本
- 验证合并对图像分类、生成和文本任务的组合泛化能力提升
- 明确不同方法对模型数量、计算资源的依赖条件
模型合并已成为低成本整合多个模型能力并提升性能的常用方法。该方法的流行推动了众多新合并技术的快速发展,但这些方法常在不同的实验设置下评估,且对模型架构、数据可用性和计算预算的假设各不相同。本文通过在统一实验环境中评估不同合并方法,系统刻画其相对优势,并精确识别每种方法的实际需求。实验聚焦于图像分类、图像生成和自然语言处理中的组合泛化能力提升,同时测量各类方法的计算开销,并研究合并模型数量扩展时的表现。结果清晰揭示了模型合并领域的现状,为未来方法的评测提供了全面且严谨的实验基准。
原文摘要 · Abstract (English)
Merging has become a widespread way to cheaply combine individual models into a single model that inherits their capabilities and attains better performance. This popularity has spurred rapid development of many new merging methods, which are typically validated in disparate experimental settings and frequently differ in the assumptions made about model architecture, data availability, and computational budget. In this work, we characterize the relative merits of different merging methods by evaluating them in a shared experimental setting and precisely identifying the practical requirements of each method. Specifically, our setting focuses on using merging for compositional generalization of capabilities in image classification, image generation, and natural language processing. Additionally, we measure the computational costs of different merging methods as well as how they perform when scaling the number of models being merged. Taken together, our results clarify the state of the field of model merging and provide a comprehensive and rigorous experimental setup to test new methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。