arXiv:2504.18884cs.CLcs.AI2025-04中稿 · the 30th Internati…被引 14

用多次中等模型推理集成,提升文本分类稳定性和准确率

A Simple Ensemble Strategy for LLM Inference: Towards More Stable Text Classification

  • 对多个中等规模LLM的推理结果做集成,替代单次大模型推理
  • 相比单次大模型推理,RMSE降低18.6%,结果更稳定准确
  • 适合需要高可靠性的文本分类任务,尤其在资源有限时

随着大语言模型(LLMs)的发展,其被广泛应用于各类任务。然而,现有研究普遍忽视了每次推理结果的波动性与可复现性问题,而人类标注通常通过多数投票解决分歧。为此,本研究引入简单的集成策略用于情感分析。实验表明,使用多个中等规模的LLM进行多次推理并集成,所得结果比仅用一次大模型推理更稳健、更准确,将均方根误差(RMSE)降低了18.6%。

原文摘要 · Abstract (English)

With the advance of large language models (LLMs), LLMs have been utilized for the various tasks. However, the issues of variability and reproducibility of results from each trial of LLMs have been largely overlooked in existing literature while actual human annotation uses majority voting to resolve disagreements among annotators. Therefore, this study introduces the straightforward ensemble strategy to a sentiment analysis using LLMs. As the results, we demonstrate that the ensemble of multiple inference using medium-sized LLMs produces more robust and accurate results than using a large model with a single attempt with reducing RMSE by 18.6%.

文本分类集成学习大模型推理稳定性提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。