arXiv:2609.03370cs.CL2026-09

测试大模型能否像人一样理解词语背后的语义框架。

FrameBench:A Language Understanding Benchmark Based on Frame Semantics

论文配图:FrameBench:A Language Understanding Benchmark Based on Frame Semantics
图 1 · 摘自论文原文
  • 基于语义框架构建多选题,考察模型对同一动词在不同语境中的理解差异。
  • 小模型表现不佳,部分大模型得分超过人类参考水平。
  • 适用于评估模型深层语义理解能力,适合关注推理与常识的科研人员。

在语义框架理论中,句子理解依赖于词汇意义与背景知识(即语义框架)的关联,使读者能隐式补充文本中未明说的信息。近年来大型语言模型在诸多下游任务中表现优异,但其是否具备人类自然理解时的隐含信息补全能力仍不明确。为此,我们提出FrameBench,一个基于语义框架的基准测试。该基准包含多项选择题,用于检验模型能否区分同一动词在不同语境下所唤起的不同语义框架。我们利用FrameNet风格资源及母语者参与的生成与验证流程,在英语和日语中构建了该数据集。对多种模型的实验表明,小模型面临挑战,而部分大模型得分已超过人类参考分数。我们已将构建的FrameBench数据集及代码发布于https://github.com/SasanoLab/FrameBench。

原文摘要 · Abstract (English)

In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at https://github.com/SasanoLab/FrameBench.

语义理解框架语义大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。