arXiv:2411.14103cs.CL2024-11被引 10

NLI任务仍能有效区分大模型能力,且与人类判断更接近。

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models

  • 用五项NLI基准测试六种规模大模型,评估其区分能力。
  • 模型准确率随训练进展提升,未达饱和,具动态判别力。
  • 大模型输出分布更贴近人类,适合衡量模型理解水平。

近期评估大语言模型(LLM)自然语言理解能力的主流方法是考察其在自然语言推理(NLI)任务上的表现。本文研究发现,尽管NLI任务很少用于当前大模型评估,但仍具有信息价值。我们在六种不同规模的模型上,基于五个跨领域的NLI基准进行测试,验证其能否区分不同规模与质量的模型,并分析准确率在训练过程中的演变。此外,我们还考察了当语句存在歧义或模糊时,模型的softmax输出分布与人类标注分布的一致性。结果表明:NLI任务能有效区分各训练阶段的模型,且未完全饱和;模型输出分布与人类分布的相似性随规模增大而提升,甚至超过两组人类群体间的相似度,说明其可作为有价值的评估指标。

原文摘要 · Abstract (English)

In the recent past, a popular way of evaluating natural language understanding (NLU), was to consider a model's ability to perform natural language inference (NLI) tasks. In this paper, we investigate if NLI tasks, that are rarely used for LLM evaluation, can still be informative for evaluating LLMs. Focusing on five different NLI benchmarks across six models of different scales, we investigate if they are able to discriminate models of different size and quality and how their accuracies develop during training. Furthermore, we investigate the extent to which the softmax distributions of models align with human distributions in cases where statements are ambiguous or vague. Overall, our results paint a positive picture for the NLI tasks: we find that they are able to discriminate well between models at various stages of training, yet are not (all) saturated. Furthermore, we find that while the similarity of model distributions with human label distributions increases with scale, it is still much higher than the similarity between two populations of humans, making it a potentially interesting statistic to consider.

NLI大模型评估语言理解模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。