构建多语言习语数据集,揭示低资源语言习语理解难题
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

- 收集高、中、低资源语言的句子与对话级习语数据
- 发现低资源语言习语理解性能显著下降,字面义比隐喻义更难识别
- 揭示当前模型依赖记忆而非推理,尤其在对话上下文中表现不足
习语表达因意义在字面与隐喻间切换,常需上下文才能准确理解,是多语言自然语言处理的重大挑战。以往研究集中于高资源语言,仅评估孤立的习语-含义匹配问题,忽视真实语篇。本文引入MIDI数据集,涵盖3种高资源、3种中资源和12种低资源语言,由母语者精心标注。与已有数据集不同,MIDI提供嵌入在句子和对话上下文中的习语,涵盖字面与隐喻两种解读。基准测试显示,习语理解在低资源语言中性能显著下降;且在所有资源层级,字面义理解均远难于隐喻义。对话上下文虽能提升性能,但无法消除这些差距。通过控制实验与对隐藏表示的干预分析,进一步区分了模型的记忆与推理能力,暴露当前模型的核心局限。
原文摘要 · Abstract (English)
Idiomatic expressions pose a major challenge for multilingual NLP because their meanings shift between figurative and literal usage, often requiring context for accurate interpretation. Prior work has focused on high-resource languages typically evaluates isolated idiom-meaning questions, overlooking realistic discourse. We introduce MIDI, a multilingual idiom dataset spanning 3 high-, 3 medium-, and 12 low-resource languages, curated by native speakers. Unlike previous datasets, MIDI provides idioms embedded in both sentence-level and conversational contexts, capturing both literal and figurative readings. Benchmarking state-of-the-art models shows that idiom comprehension degrades in low-resource languages and that, in all resource tiers, literal interpretations are substantially harder than figurative ones. Conversational context improves performance but does not eliminate these disparities. Through controlled tests and interventions on hidden representations, we further separate memorization from reasoning, exposing core limitations of current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。