arXiv:2603.18678cs.SDcs.CL2026-03

首个评测大模型理解语音双关的基准,揭示幽默感知关键难题

Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models

  • 构建4434个语音双关样本的三阶段标注数据集
  • 10个主流大音频语言模型在定位与释义上表现差距显著
  • 发现位置偏差和语义误判问题,推动幽默感知模型发展

双关语是利用多义性和发音歧义制造幽默的典型语言现象,对自然语言理解构成独特挑战。在双关研究中,语音在人类交流中占据核心地位,但针对口语双关的数据集和系统性资源仍十分稀缺,导致该关键模态长期未受重视。本文提出APUN-Bench,首个专用于评估大音频语言模型(LALMs)在语音双关理解能力上的基准。该基准包含4,434个音频样本,经三阶段标注:双关识别、双关词定位与双关含义推断。我们系统评估了10个前沿LALMs,发现其在识别、定位与解释语音双关方面存在显著性能差距。分析揭示了双关定位中的位置偏差及语义推断错误等关键挑战,为发展具备幽默感知能力的音频智能提供了可操作的改进方向。

原文摘要 · Abstract (English)

Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges for natural language understanding. Within pun research, audio plays a central role in human communication except text and images, while datasets and systematic resources for spoken puns remain scarce, leaving this crucial modality largely underexplored. In this paper, we present APUN-Bench, the first benchmark dedicated to evaluating large audio language models (LALMs) on audio pun understanding. Our benchmark contains 4,434 audio samples annotated across three stages: pun recognition, pun word location and pun meaning inference. We conduct a deep analysis of APUN-Bench by systematically evaluating 10 state-of-the-art LALMs, uncovering substantial performance gaps in recognizing, localizing, and interpreting audio puns. This analysis reveals key challenges, such as positional biases in audio pun location and error cases in meaning inference, offering actionable insights for advancing humour-aware audio intelligence.

语音理解双关语大模型评测幽默感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。