多语言训练提升谚语隐喻识别,文化特异性形式受益最大。
Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs

- 用七种语言742个谚语概念构建多维标注框架
- 约50%多语言数据即可达近优性能,文化特异性形式增益最显著
- 指令微调大模型更依赖道德与文化类隐喻特征
尽管多语言隐喻识别已有研究,但跨语言训练数据的贡献仍需厘清。本文基于742个谚语概念、6,787个跨语言翻译实例,在七种语言中评估五种模型,涵盖多语言编码器与指令微调大模型。提出包含四类互补隐喻形式(隐喻、道德/劝诫、因果、文化特异性)的多维标注框架。结果表明,约50%的多语言训练数据即可实现接近最优的识别性能;融合多种隐喻形式效果最佳。值得注意的是,最不常见的文化特异性形式在多语言监督下表现提升最大。道德/劝诫与文化特异性形式对指令微调大模型性能贡献最显著。研究呼吁从以隐喻为中心的分类转向以概念为核心的多维度隐喻建模。
原文摘要 · Abstract (English)
Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We examine this question using 742 proverb concepts across 6,787 translated instances in seven languages. We evaluate five models, including multilingual encoders and instruction tuned LLMs, under progressively increasing levels of multilingual supervision. Moreover, we introduce a multidimensional annotation framework for proverbs that characterizes them through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture Specific. Our findings show that approximately 50% of the translated multilingual training data is sufficient to achieve near-optimal figurative language identification performance. We further show that combining diverse figurative forms yields the strongest overall performance. A notable finding is that the least frequent figurative form, Culture Specific, exhibits the largest performance gains under multilingual supervision. Furthermore, the Moral/Advisory and Culture Specific forms contribute most to the performance of instruction-tuned LLMs on figurative language identification. These findings motivate multilingual figurative language identification to move beyond metaphor-centric taxonomies toward concept level multidimensional frameworks that explicitly model complementary forms of figurative meaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。