arXiv:2604.16593cs.CL2026-04

构建语义理解测评基准,检验大模型对复杂短语的推理能力。

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

论文配图:Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models
图 1 · 摘自论文原文
  • 整合多词表达数据,统一为可评测的语义任务集。
  • 发现不同模型在语义推理任务上表现差异显著,尤其在习语理解上短板明显。
  • 适合研究语言模型语义理解与推理能力的研究者使用。

我们提出SemanticQA,一个用于评估语言模型(LMs)在语义短语处理任务中表现的评测套件。该基准整合了现有的多词表达(MwE)资源,并重新组织为统一测试平台,涵盖一般词汇现象(如搭配)、习语、名词复合词和动词构式三类细粒度类别。通过SemanticQA,我们评估了多种架构与规模的语言模型在提取、分类与解释任务以及序列任务组合中的表现。结果揭示出显著的性能差异,尤其在需要语义推理的任务中,凸显了不同模型在推理效率与语义理解上的差距,为提升模型对非平凡语义短语的理解能力提供了洞见。SemanticQA的评测工具与数据可在https://github.com/jacklanda/SemanticQA获取。

原文摘要 · Abstract (English)

We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates existing multiword expression (MwE) resources and reorganizes them into a unified testbed. It covers both general lexical phenomena, such as lexical collocations, and three fine-grained categories: idiomatic expressions, noun compounds, and verbal constructions. Through SemanticQA, we assess LMs of diverse architectures and scales in extraction, classification, and interpretation tasks, as well as sequential task compositions. We reveal substantial performance variation, particularly on tasks requiring semantic reasoning, highlighting differences in reasoning efficacy and semantic understanding of LMs, providing insights for pushing LMs with stronger comprehension on non-trivial semantic phrases. The evaluation harness and data of SemanticQA are available at https://github.com/jacklanda/SemanticQA.

语义理解语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。