提出文化感知的嵌套表示框架,区分阿拉伯语隐喻中的词汇、文化与隐喻信息。
CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations

- 分层嵌套设计,按文化理论逐步构建词汇-文化-隐喻表示空间。
- 几何度量隐喻性,配对监督下检测准确率AUC达0.84,正确识别82.6%隐喻对。
- 无需训练,适合研究阿拉伯语文化隐喻与语言模型表征的学者。
阿拉伯语隐喻是根植于文化的表意机制,编码文化知识以塑造理解。然而现有阿拉伯语模型常将词汇、文化与隐喻信息混同于单一表征空间,造成“语义模糊”。本文提出CAMMAR(文化感知嵌套式阿拉伯语隐喻表征框架),通过阶段性语义课程,将意义组织为嵌套的词汇、文化与隐喻子空间。该设计基于阿尔-朱尔贾尼的nazum理论,将修辞意义建模为先前语义关系的组合,生成无需训练的几何隐喻度量——基于词汇与隐喻表示间的距离。在新构建的跨度标注阿拉伯语隐喻数据集上评估,配对监督下几何读出的隐喻检测优于随机水平(AUC最高0.84),同一词的隐喻对在82.6%情况下优于其字面对应;而仅靠无监督领域对比时则表现如随机,清晰区分出有监督可学与无监督无法涌现的两种情形。受控消融实验显示,以词根锚定词汇层带来微小但稳定的提升,此效应在直接探针中缺失,说明其作为测量基准的质量。数据集、文化概念库与代码将在录用后公开。
原文摘要 · Abstract (English)
Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet current Arabic language models typically collapse lexical, cultural, and metaphorical information into a single representational space, a phenomenon we term "semantic smearing". We introduce CAMMAR (Culture-Aware Matryoshka for Metaphorical Arabic Representations), a representation learning framework that organizes meaning into nested lexical, cultural, and metaphorical embedding subspaces through a staged semantic curriculum. The design implements compositional principles of Al-Jurjani's theory of nazum, modeling figurative meaning as compositionally grounded in prior semantic relations, and yields a training-free geometric measure of metaphoricity based on the distance between lexical and metaphorical representations. Evaluated on a new span-annotated Arabic metaphor set as word-matched figurative/literal pairs, the geometric readout detects metaphor well above chance when the inter-layer geometry is shaped by paired supervision (AUC up to 0.84; figurative outscores its literal counterpart for the same word in 82.6\% of pairs), but sits at chance under an unsupervised domain contrast alone, a clean separation between a legible-under-supervision regime and a non-emergent one. A controlled ablation shows that grounding the lexical layer in morphological roots gives a small but consistent gain, an effect absent from direct probing that reflects the layer's quality as a measurement anchor. We will release the datasets, cultural concept inventory, and code upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。