arXiv:2608.21087cs.CL2026-08

用语义向量测量双关语的歧义距离,发现对称性或为好笑的关键。

Jokes Aside: Measuring the Semantic Distance of Double Meanings

  • 用嵌入向量重测双关笑话的五个指标,新增对称性度量
  • 对称性越高笑话越有趣,但整体模型预测准确率不足60%
  • 适合研究幽默机制或语义模糊性的学者参考

大语言模型为计算幽默研究提供了新工具,尤其在笑话和双关语自动生成方面。基于2013年提出的「我喜欢我的X就像我喜欢我的Y,Z」结构,本文重新评估了其中四个假设:Z与X、Y的频繁关联、稀有性、歧义性以及X与Y之间的语义距离。借鉴Winters等(2019)的五项指标,本研究采用开放AI text-embedding-3-small和MiniLM all-MiniLM-L6-v2两个模型,在JokeJudger、Expunations和rJokes三个数据集上提取嵌入向量。其中,Expunations和rJokes通过添加双关核心词的双义句对进行扩展。结果表明,基于新指标训练的模型在JokeJudger上最高仅达57.1%准确率,低于61.5%基线;在另两个数据集上表现更差。然而,新引入的对称性(即Z与X、Y的语义接近程度)与高评分笑话存在一致关联,提示其可能是幽默的必要而非充分条件。

原文摘要 · Abstract (English)

Large language models have significantly enriched the toolkit for computational humor research, particularly in the automated generation of jokes and puns. A key innovation, contextual embedding vectors, offers new opportunities to revisit and refine earlier hypotheses. Notably, Petrovic and Matthews (2013) proposed a joke generation model based on the scheme "I like my X like I like my Y, Z" (e.g. "I like my ice like I like my dreams, crushed"). They suggested that joke hilarity increases with: a) frequent association of Z with X and Y, b) rarity of Z, c) ambiguity of Z, and d) meaning distance between X and Y. Building on this, Winters et al. (2019) proposed a set of metrics, based on Google Ngrams and Word2Vector. In this work, three out of their five metrics are revisited with word embeddings: obviousness, compatibility, and comparison. Another measure, symmetry, defined as closeness of Z to both X and Y, is introduced here for the first time. Two models were used to collect the embedding vectors (OpenAI text-embedding-3-small and MiniLM all-MiniLM-L6-v2) on three datasets: JokeJudger, Expunations, and rJokes. The last two datasets, Expunations, and rJokes, were expanded by adding paired sentences that captured the ambiguous expression at the core of each joke in its two different meanings. Results revealed that models trained on the proposed metrics performed poorly in predicting humor ratings: on JokeJudger, the best model achieved 57.1% accuracy, below the 61.5% baseline, while performance on Expunations and rJokes was even lower. Nevertheless, the symmetry metric seems consistently associated with higher-rated jokes, suggesting it may capture a necessary -though not sufficient- property of humor.

双关语语义距离幽默生成嵌入向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。