arXiv:2502.15810cs.CLcs.AI2025-02被引 2

评测大模型零样本理解常识的能力,发现超大规模模型在判断常识对错上接近人类水平。

Zero-Shot Commonsense Validation and Reasoning with Large Language Models: An Evaluation on SemEval-2020 Task 4 Dataset

  • 用零样本提示测试多个大模型在常识验证与解释任务上的表现
  • LLaMA3-70B在常识验证任务中达到98.40%准确率,接近人类水平
  • 模型能识别不合理陈述但难选最相关解释,暴露推理缺陷

本研究评估大型语言模型在SemEval-2020 Task 4数据集上的常识验证与解释能力。采用零样本提示方法,测试包括LLaMA3-70B、Gemma2-9B和Mixtral-8x7B在内的多个模型。任务分为两部分:任务A(常识验证)判断语句是否符合常识;任务B(常识解释)识别不合理语句背后的推理原因。基于准确率评估结果表明,更大规模模型性能优于先前模型,在任务A中表现接近人类,其中LLaMA3-70B达到最高准确率98.40%;但在任务B中仅达93.40%,落后于已有微调模型。尽管模型能有效识别不合理的陈述,却难以选出最相关的解释,暴露出因果与推断推理的局限性。

原文摘要 · Abstract (English)

This study evaluates the performance of Large Language Models (LLMs) on SemEval-2020 Task 4 dataset, focusing on commonsense validation and explanation. Our methodology involves evaluating multiple LLMs, including LLaMA3-70B, Gemma2-9B, and Mixtral-8x7B, using zero-shot prompting techniques. The models are tested on two tasks: Task A (Commonsense Validation), where models determine whether a statement aligns with commonsense knowledge, and Task B (Commonsense Explanation), where models identify the reasoning behind implausible statements. Performance is assessed based on accuracy, and results are compared to fine-tuned transformer-based models. The results indicate that larger models outperform previous models and perform closely to human evaluation for Task A, with LLaMA3-70B achieving the highest accuracy of 98.40% in Task A whereas, lagging behind previous models with 93.40% in Task B. However, while models effectively identify implausible statements, they face challenges in selecting the most relevant explanation, highlighting limitations in causal and inferential reasoning.

大模型常识推理零样本评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。