小模型答同一题反复出错,一致性仅50%-80%。
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
- 测试小模型在相同题目上重复作答的一致性
- 2B-8B模型一致性普遍50%-80%,温度越低越稳定
- 适合关注模型可靠性与评测设计的研究者
本研究考察了小型大语言模型(2B-8B参数)在多次重复回答同一问题时的表现一致性。我们对开源小模型在MMLU-Redux和MedQA两个多选基准上进行10次重复测试,分析不同推理温度、小模型(2B-8B)与中等模型(50B-80B)、微调版与基础版模型的差异。结果表明,小模型在低温度下的一致性通常介于50%-80%之间;一致答案的准确率与整体准确率呈合理相关性。中等规模模型表现出更高的一致性水平。研究还提出新的分析与可视化工具,以支持对答案一致性与准确率权衡的评估。
原文摘要 · Abstract (English)
This work explores the consistency of small LLMs (2B-8B parameters) in answering multiple times the same question. We present a study on known, open-source LLMs responding to 10 repetitions of questions from the multiple-choice benchmarks MMLU-Redux and MedQA, considering different inference temperatures, small vs. medium models (50B-80B), finetuned vs. base models, and other parameters. We also look into the effects of requiring multi-trial answer consistency on accuracy and the trade-offs involved in deciding which model best provides both of them. To support those studies, we propose some new analytical and graphical tools. Results show that the number of questions which can be answered consistently vary considerably among models but are typically in the 50%-80% range for small models at low inference temperatures. Also, accuracy among consistent answers seems to reasonably correlate with overall accuracy. Results for medium-sized models seem to indicate much higher levels of answer consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。