专门测试大模型对句子级否定理解能力的新基准。
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
- 设计对比标准否定与多种结构变体的句子对
- 包含人工标注的否定句对和多选题数据集
- 适合评估模型在深层语义理解中的否定处理能力
否定是语言学中的基本现象,对大型语言模型(LLMs)构成持续挑战,尤其是在需要深度语义理解的任务中。现有基准通常将否定视为更广泛任务中的次要细节,如自然语言推理,因此缺乏专门用于评估否定理解的基准。本文提出Thunder-NUBench,一个专为评估大模型句子级否定理解而设计的新基准。该基准不仅识别表面线索,还对比标准否定与结构多样化的替代形式,如局部否定、矛盾句和改写句。基准包含人工精心构建的句子-否定对和多选题数据集,支持对模型否定理解能力的全面评估。
原文摘要 · Abstract (English)
Negation is a fundamental linguistic phenomenon that poses ongoing challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Current benchmarks often treat negation as a minor detail within broader tasks, such as natural language inference. Consequently, there is a lack of benchmarks specifically designed to evaluate comprehension of negation. In this work, we introduce Thunder-NUBench, a novel benchmark explicitly created to assess sentence-level understanding of negation in LLMs. Thunder-NUBench goes beyond merely identifying surface-level cues by contrasting standard negation with structurally diverse alternatives, such as local negation, contradiction, and paraphrase. This benchmark includes manually curated sentence-negation pairs and a multiple-choice dataset, allowing for a comprehensive evaluation of models' understanding of negation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。