语言模型对隐喻新颖性的预测能力有限,且表现随数据类型不同而异。
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets
- 用全句上下文的填空式可预测性衡量隐喻词的新颖性
- 在真实语料中相关性随模型增大而下降,在合成数据中上升
- 揭示了语言模型创造力评估的局限性,适合对隐喻认知感兴趣的研究者
隐喻理解涉及复杂的语义过程和语言创造性,是研究语言模型(LMs)的有趣任务。本研究探讨了可预测性度量 surprisal 是否与不同数据集中的隐喻新颖性标注相关。我们使用16种因果语言模型变体,分析了语料库和合成隐喻数据集中隐喻词的 surprisal。提出一种基于完整句子上下文的填空式 surprisal 方法。结果显示,语言模型的 surprisal 与隐喻新颖性评分/标签存在显著中等程度的相关性。进一步发现,两种数据类型呈现相反的缩放模式:在语料库数据中,相关性随模型规模增大而减弱(反向缩放效应),而在合成数据中则增强(质量-功率假说)。结论认为,尽管 surprisal 能部分解释隐喻新颖性标注,但作为语言创造力的度量仍具局限性。代码与数据已公开:https://github.com/OmarMomen14/surprisal-metaphor-novelty
原文摘要 · Abstract (English)
Novel metaphor comprehension involves complex semantic processes and linguistic creativity, making it an interesting task for studying language models (LMs). This study investigates whether surprisal, a probabilistic measure of predictability in LMs, correlates with annotations of metaphor novelty in different datasets. We analyse the surprisal of metaphoric words in corpus-based and synthetic metaphor datasets using 16 causal LM variants. We propose a cloze-style surprisal method that conditions on full-sentence context. Results show that LM surprisal yields significant moderate correlations with scores/labels of metaphor novelty. We further identify divergent scaling patterns: on corpus-based data, correlation strength decreases with model size (inverse scaling effect), whereas on synthetic data it increases (quality-power hypothesis). We conclude that while surprisal can partially account for annotations of metaphor novelty, it remains limited as a metric of linguistic creativity. Code and data are publicly available: https://github.com/OmarMomen14/surprisal-metaphor-novelty
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。