测试大模型对深层荒诞语的理解能力,发现它们常误判为普通废话。
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
- 构建跨语言的'深度荒诞'语料库,含1200+经专家验证的例句。
- 模型在分类、生成和推理任务中普遍误判隐含意义,准确率偏低。
- 适合研究语言理解深度、认知推理与大模型局限性的学者。
我们提出Drivelology,一种表现为'有深度的荒诞'的语言现象——语法通顺但语用悖论、情感浓烈或修辞颠覆。这类表达表面看似无意义,实则需上下文推断、道德判断或情感解读。尽管当前大语言模型(LLMs)在众多自然语言处理任务中表现优异,却难以理解此类文本的多层语义。为此,我们构建了一个包含1200多个例句的基准数据集,覆盖英语、中文、西班牙语、法语、日语和韩语,每个样本均经过多轮专家评审与争议仲裁以确保其符合Drivelological特征。利用该数据集,我们评估了多种LLMs在分类、生成与推理任务中的表现。结果表明,模型常将深度荒诞误认为浅层废话,生成逻辑混乱的解释,或完全忽略隐含修辞功能。这些发现揭示了大模型在语用理解上的深层表征缺陷,挑战了统计流畅即具认知理解的假设。我们已开源数据集与代码,以推动对超越表面连贯性的语言深度建模研究。
原文摘要 · Abstract (English)
We introduce Drivelology, a unique linguistic phenomenon characterised as "nonsense with depth" - utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or rhetorically subversive. While such expressions may resemble surface-level nonsense, they encode implicit meaning requiring contextual inference, moral reasoning, or emotional interpretation. We find that current large language models (LLMs), despite excelling at many natural language processing (NLP) tasks, consistently fail to grasp the layered semantics of Drivelological text. To investigate this, we construct a benchmark dataset of over 1,200+ meticulously curated and diverse examples across English, Mandarin, Spanish, French, Japanese, and Korean. Each example underwent careful expert review to verify its Drivelological characteristics, involving multiple rounds of discussion and adjudication to address disagreements. Using this dataset, we evaluate a range of LLMs on classification, generation, and reasoning tasks. Our results reveal clear limitations of LLMs: models often confuse Drivelology with shallow nonsense, produce incoherent justifications, or miss implied rhetorical functions altogether. These findings highlight a deep representational gap in LLMs' pragmatic understanding and challenge the assumption that statistical fluency implies cognitive comprehension. We release our dataset and code to facilitate further research in modelling linguistic depth beyond surface-level coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。