arXiv:2510.21575cs.CLcs.AI2025-10

首个斯洛文尼亚语语用理解基准,助力大模型读懂隐含意图。

From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene

  • 构建斯洛文尼亚语语用理解评测集,含405道多选题。
  • 现有模型在文化特定隐喻理解上仍存在明显短板。
  • 适合研究多语言语用、跨文化语言理解的学者使用。

大型语言模型的能力不断提升,已在以往被认为极难的基准测试中表现优异。随着能力增强,亟需更复杂的评估体系,超越表层语言能力,深入考察语用理解——即结合语境、语言和文化规范理解情境意义。为此,我们提出了SloPragEval与SloPragMega,这是首个针对斯洛文尼亚语的语用理解基准,共包含405道多选题。我们讨论了翻译难度,描述了建立人类基线的征集活动,并报告了对大模型的初步评估结果。结果显示,当前模型在理解复杂语言方面已有显著进步,但在推断非字面表达中说话者的隐含意图,特别是文化特异性表达时仍易失败。同时观察到专有模型与开源模型之间存在显著差距。最后我们强调,面向语用理解和文化知识的评测必须精心设计,最好基于母语数据构建,并通过人类回答进行验证。

原文摘要 · Abstract (English)

Large language models are demonstrating increasing capabilities, excelling at benchmarks once considered very difficult. As their capabilities grow, there is a need for more challenging evaluations that go beyond surface-level linguistic competence. Namely, language competence involves not only syntax and semantics but also pragmatics, i.e., understanding situational meaning as shaped by context as well as linguistic and cultural norms. To contribute to this line of research, we introduce SloPragEval and SloPragMega, the first pragmatics understanding benchmarks for Slovene that contain altogether 405 multiple-choice questions. We discuss the difficulties of translation, describe the campaign to establish a human baseline, and report pilot evaluations with LLMs. Our results indicate that current models have greatly improved in understanding nuanced language but may still fail to infer implied speaker meaning in non-literal utterances, especially those that are culture-specific. We also observe a significant gap between proprietary and open-source models. Finally, we argue that benchmarks targeting nuanced language understanding and knowledge of the target culture must be designed with care, preferably constructed from native data, and validated with human responses.

语用理解多语言评测基准文化差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。