arXiv:2412.11763cs.CL2024-12

测试大模型在地理常识推理上的跨语言能力差距

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs

  • 基于油管问答视频构建带掩码实体的英文测验集
  • 7个大模型在零样本下平均准确率仅41.2%
  • 适合评估多语言常识推理能力的基准研究

大语言模型(LLMs)的兴起催生了超越传统设置的先进评测体系需求。为此,我们提出QUENCH——一个从油管问答视频中人工整理并转录的文本型英文测验基准。QUENCH包含被掩码的实体和推理链,要求模型通过生成进行预测。该基准融合地理语境与常识推理,通过零样本、开放域测验设定评估模型的世界知识与推断能力。我们在7个大语言模型上使用4种指标进行了广泛评估,分析了模型规模、提示风格、地理语境及标注推理链生成的影响。最终通过误差分析揭示了模型易犯的错误类型。

原文摘要 · Abstract (English)

The rise of large language models (LLMs) has created a need for advanced benchmarking systems beyond traditional setups. To this end, we introduce QUENCH, a novel text-based English Quizzing Benchmark manually curated and transcribed from YouTube quiz videos. QUENCH possesses masked entities and rationales for the LLMs to predict via generation. At the intersection of geographical context and common sense reasoning, QUENCH helps assess world knowledge and deduction capabilities of LLMs via a zero-shot, open-domain quizzing setup. We perform an extensive evaluation on 7 LLMs and 4 metrics, investigating the influence of model size, prompting style, geographical context, and gold-labeled rationale generation. The benchmarking concludes with an error analysis to which the LLMs are prone.

大模型评测常识推理跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。