通过相对过拟合机制,实现可控的模型集成,突破大模型性能瓶颈。
Relative Overfitting and Accept-Reject Framework
- 基于模型细粒度分割与相对过拟合概念,构建可控制的集成框架
- 在语言建模、长文本任务等多基准上实现稳定性能提升
- 为大模型集成提供理论支撑,适用于NLP、CV及科学计算领域
当前大语言模型(LLMs)的扩展面临显著挑战。模型组装被视为突破性能瓶颈的有前景方案。然而,现有集成方法主要依赖统计期望:在大规模样本上组合多个模型可带来性能提升。本文提出一种从随机、依赖样本的方法转向规则化、可控方法的集成框架,基于细粒度模型分割,规范模型分割方式以确保性能提升,揭示提升幅度随模型选择的变化规律,并确定其理论上限。为此,我们引入‘相对过拟合’概念,由各子模型间的性能差异推导而来,建立起集成结果与模型内在属性之间的桥梁。我们在NLP领域详细阐述该框架的模式,并简要说明其在计算机视觉(CV)和科学人工智能中的可拓展性。实验使用自建与主流预训练模型,在语言建模、长上下文任务及问答(QA)等多个基准上验证,结果表明所提集成规则普遍有效,并在某些实验场景中提供了严格证明。该框架为理解集成理论提供了新视角,为解决大模型性能瓶颈提供了系统性方法。
原文摘要 · Abstract (English)
The scaling of Large Language Models (LLMs) currently faces significant challenges. Model assembly is widely considered a promising solution to break through these performance bottlenecks. However, current ensembling methods are primarily guided by the statistical expectation that combining multiple models over large samples will lead to performance gains. We propose an ensemble framework that transitions from such stochastic, sample-dependent methods to a regular, controllable approach based on fine-grained model segmentation. This regularity governs how models are segmented to ensure performance improvement, how the magnitude of this improvement varies with model selection, and what factors determine its theoretical maximum. To formalize this pattern, we introduce the concept of'relative overfitting,' which is derived from the performance discrepancies between constituent models and builds a bridge between ensemble outcomes and the inherent attributes of these models. We detail the patterns of this framework within the domain of NLP and briefly describe its extensibility to other fields, such as computer vision (CV) and AI for science. Our approach was validated using both custom-built and pre-trained mainstream models across diverse benchmarks, including language modeling, long-context tasks, and question-answering (QA). The results indicate that the ensemble rules we proposed are generally effective and that we provide a rigorous proof of these rules in certain experimental scenarios. The proposed framework offers a new perspective for understanding ensemble theory and provides a systematic approach to addressing the performance bottlenecks of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。