研究大模型大小如何影响知识图谱任务表现,发现小模型在某些场景更划算。
How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance
- 基于26个主流大模型,测试其在知识图谱任务中的性能随规模变化规律。
- 多数任务中模型越大越强,但部分出现性能饱和,小模型已足够。
- 同系列模型中大不一定更好,建议对比相邻大小模型选最优性价比。
在利用大语言模型(LLMs)支持知识图谱工程(KGE)任务时,模型大小是首要考虑因素。根据缩放定律,模型越大通常能力越强。然而实际应用中资源成本同样重要,需权衡性能与开销。LLM-KG-Bench框架可评估LLMs在理解与生成知识图谱及其查询方面的能力。基于该框架运行生成的涵盖26个前沿大模型的数据集,我们分析了特定于KGE任务的模型规模缩放规律。研究发现,尽管多数情况下模型规模越大性能越高,但存在局部平台期或天花板效应——即相邻模型间性能提升不明显。此时较小模型更具成本效益。此外,在同一模型家族内,有时更大模型表现反而更差,此类现象虽局部存在,但仍建议对同族模型进行相邻规模对比测试,以实现最优选择。
原文摘要 · Abstract (English)
When using Large Language Models (LLMs) to support Knowledge Graph Engineering (KGE), one of the first indications when searching for an appropriate model is its size. According to the scaling laws, larger models typically show higher capabilities. However, in practice, resource costs are also an important factor and thus it makes sense to consider the ratio between model performance and costs. The LLM-KG-Bench framework enables the comparison of LLMs in the context of KGE tasks and assesses their capabilities of understanding and producing KGs and KG queries. Based on a dataset created in an LLM-KG-Bench run covering 26 open state-of-the-art LLMs, we explore the model size scaling laws specific to KGE tasks. In our analyses, we assess how benchmark scores evolve between different model size categories. Additionally, we inspect how the general score development of single models and families of models correlates to their size. Our analyses revealed that, with a few exceptions, the model size scaling laws generally also apply to the selected KGE tasks. However, in some cases, plateau or ceiling effects occurred, i.e., the task performance did not change much between a model and the next larger model. In these cases, smaller models could be considered to achieve high cost-effectiveness. Regarding models of the same family, sometimes larger models performed worse than smaller models of the same family. These effects occurred only locally. Hence it is advisable to additionally test the next smallest and largest model of the same family.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。