选好例子+调高温度,让大模型少犯错
Stable LLM Ensemble: Interaction between Example Representativeness and Diversity
- 用代表性例子替代随机选例,提升预测质量
- 高温采样使模型表现优于随机选例7.6%(宏F1)
- 适合追求高效准确的模型部署与系统设计
大型语言模型在多个领域已取得显著成果。然而,单样本提示下的预测准确性和鲁棒性仍高度依赖于示例选择和集成成员间的多样性。本研究系统考察了示例代表性(单样本策略)与输出多样性(采样温度)对语言模型集成性能的影响。对比了两种单样本策略:基于质心的代表性示例(新提出)与随机采样示例(基线),并调节采样温度。结果表明,采用更高温度设置的新方法相比随机选择,在宏平均F1上提升7.6%,均方根误差降低10.5%。此外,该方法性能超过五样本提示,宏F1提升21.1%,均方根误差下降24.0%。研究证实,结合代表性示例选择与适度提高温度,可为集成提供恰到好处的多样性。本工作强调了示例选择与可控多样性在构建高效单样本语言模型集成中的实际意义。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable results in wide range of domains. However, the accuracy and robustness of one-shot LLM predictions remain highly sensitive to the examples and the diversity among ensemble members. This study systematically investigates the effects of example representativeness (one-shot strategy) and output diversity (sampling temperature) on LLM ensemble performance. Two one-shot strategies are compared: centroid-based representative examples (proposed) and randomly sampled examples (baseline) and sampling temperature also is varied. The proposed approach with higher temperature setting significantly outperforms random selection by +7.6% (macro-F1) and -10.5% (RMSE). Furthermore, the proposed model exceeds 5-shot prompting by +21.1% (macro-F1) and -24.0% (RMSE). Our findings demonstrate that combining representative example selection with increased temperature provides the appropriate level of diversity to the ensemble. This work highlights the practical importance of both example selection and controlled diversity in designing effective one-shot LLM ensembles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。