跨语言委婉语迁移中,语义重叠不等于有效传递,方向性差异显著。
When Semantic Overlap Is Not Enough: Cross-Lingual Euphemism Transfer Between Turkish and English
- 按语义与语用对齐将双语委婉语分为重叠与非重叠类
- 土耳其语到英语迁移时,即使语义重叠也常性能下降
- 非重叠类训练反而在部分场景提升效果,受标签分布影响
委婉语替代敏感表达,依赖文化与语用背景,跨语言建模困难。本文研究跨语言等价性对多语言委婉语检测迁移的影响。将土耳其语和英语中的潜在委婉语(PETs)根据功能、语用和语义对齐分为重叠(OPETs)与非重叠(NOPETs)两类。结果表明存在迁移不对称性:语义重叠不足以保证正向迁移,尤其在低资源的土耳其语→英语方向,即使重叠委婉语性能也可能下降;在某些情况下,基于非重叠类的训练反而提升性能。标签分布差异可解释这些反直觉现象。类别级分析显示迁移可能受领域特定对齐影响,但数据稀疏限制了证据强度。
原文摘要 · Abstract (English)
Euphemisms substitute socially sensitive expressions, often softening or reframing meaning, and their reliance on cultural and pragmatic context complicates modeling across languages. In this study, we investigate how cross-lingual equivalence influences transfer in multilingual euphemism detection. We categorize Potentially Euphemistic Terms (PETs) in Turkish and English into Overlapping (OPETs) and Non-Overlapping (NOPETs) subsets based on their functional, pragmatic, and semantic alignment. Our findings reveal a transfer asymmetry: semantic overlap is insufficient to guarantee positive transfer, particularly in low-resource Turkish-to-English direction, where performance can degrade even for overlapping euphemisms, and in some cases, improve under NOPET-based training. Differences in label distribution help explain these counterintuitive results. Category-level analysis suggests that transfer may be influenced by domain-specific alignment, though evidence is limited by sparsity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。