arXiv:2603.11687cs.CLcs.AI2026-03中稿 · LREC 2026

用词典定义自动生成跨语言语义评测集,省时省力还准确。

SemBench: A Universal Semantic Framework for LLM Evaluation

  • 仅靠词典释义和句向量模型自动构建评测数据
  • 3种语言、多种模型测试,结果与传统基准高度一致
  • 少量样本即可稳定评估,适合多语言研究者使用

近年来自然语言处理的发展主要由大语言模型(LLMs)推动,这些模型展现出强大的生成与推理能力。然而,评估其真实语义理解能力仍是长期挑战。传统基准如词义上下文(WiC)虽有效,但构建成本高,且多限于高资源语言。本文提出SemBench,一个仅依赖词典义项定义与句编码器的自动合成评测框架,无需人工标注例句,具备可扩展性和语言无关性。我们在英语、西班牙语和巴斯克语三种语言(覆盖不同资源水平)上评估了多种LLMs,结果表明SemBench得出的排名与标准WiC数据集高度相关。此外,分析显示只需少量样本即可获得稳定可靠的排名。总体而言,SemBench为大语言模型的跨语言语义理解提供了一个轻量、灵活且高效的数据评估方案。

原文摘要 · Abstract (English)

Recent progress in Natural Language Processing (NLP) has been driven by the emergence of Large Language Models (LLMs), which exhibit remarkable generative and reasoning capabilities. However, despite their success, evaluating the true semantic understanding of these models remains a persistent challenge. Traditional benchmarks such as Word-in-Context (WiC) effectively probe this capability, but their creation is resource-intensive and often limited to high-resource languages. In this paper, we introduce SemBench, a framework for automatically generating synthetic benchmarks that assess the semantic competence of LLMs using only dictionary sense definitions and a sentence encoder. This approach eliminates the need for curated example sentences, making it both scalable and language-independent. We evaluate SemBench in three languages (English, Spanish, and Basque) spanning different levels of linguistic resources, and across a wide range of LLMs. Our results show that rankings derived from SemBench strongly correlate with those obtained from standard WiC datasets. Furthermore, our analysis demonstrates that only a small number of examples is required to achieve stable and meaningful rankings. Overall, SemBench provides a lightweight, adaptable, and data-efficient framework for cross-lingual evaluation of semantic understanding in LLMs.

语义评估大模型评测跨语言自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。