arXiv:2504.17685cs.CLcs.AI2025-04

用小模型集成实现大模型级匹配准确率

Ensemble Bayesian Inference: Leveraging Small Language Models to Achieve LLM-level Accuracy in Profile Matching Tasks

  • 通过贝叶斯融合多个小模型判断结果,突破单个模型上限
  • 在日英双语任务中达到与大模型相当的准确率
  • 即使加入表现差的模型也能提升整体性能,适合资源有限场景

本研究探索小型语言模型(SLM)集成在实现与专有大语言模型(LLM)相当的准确率方面的潜力。我们提出一种新方法——集成贝叶斯推断(EBI),通过贝叶斯估计融合多个SLM的判断,使整体性能超越单个模型的局限。在多种任务(包括能力评估和日语、英语的消费者画像分析)上的实验表明,该方法有效。值得注意的是,我们分析了将负收益模型纳入集成反而提升整体性能的案例,并考察了方法在不同语言中的有效性。这些发现为在计算资源受限条件下构建高性能AI系统提供了新路径,也展示了对低性能模型的有效利用价值。基于现有大型语言模型评估、集成学习及开源模型应用的研究,本文讨论了方法的创新性与意义。

原文摘要 · Abstract (English)

This study explores the potential of small language model(SLM) ensembles to achieve accuracy comparable to proprietary large language models (LLMs). We propose Ensemble Bayesian Inference (EBI), a novel approach that applies Bayesian estimation to combine judgments from multiple SLMs, allowing them to exceed the performance limitations of individual models. Our experiments on diverse tasks(aptitude assessments and consumer profile analysis in both Japanese and English) demonstrate EBI's effectiveness. Notably, we analyze cases where incorporating models with negative Lift values into ensembles improves overall performance, and we examine the method's efficacy across different languages. These findings suggest new possibilities for constructing high-performance AI systems with limited computational resources and for effectively utilizing models with individually lower performance. Building on existing research on LLM performance evaluation, ensemble methods, and open-source LLM utilization, we discuss the novelty and significance of our approach.

小模型集成贝叶斯推断精准匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。