构建多语言多模态习语数据集,提升AI对文化隐喻的理解能力
When Meaning Isn't Literal: Exploring Idiomatic Meaning Across Languages and Modalities
- 构建3533个印地语、孟加拉语、泰语习语的多模态数据集
- 发现大模型在习语理解上普遍存在隐喻误读问题
- 提出提示引导式解释框架,通过反馈迭代提升推理准确率
习语推理与隐喻和文化深度关联,但当前语言模型进展偏向表面词汇线索。例如,孟加拉语习语“angur fol tok”(葡萄酸)表达的是因无法获得而合理化,但朴素模型却仅关注狐狸吃葡萄的字面意象。为弥补这一缺陷,我们构建了“Mediom”,一个包含3,533个印地语、孟加拉语、泰语习语的多语言多模态语料库,每个习语均配有标准解释、跨语言翻译及对齐的图文表示。我们在Mediom上评估大型语言模型(文本推理)和视觉语言模型(形象歧义消解),揭示了其在隐喻理解上的系统性失败。为此,我们提出“HIDE”——一种基于提示的习语解释框架,通过错误反馈检索与定向诊断提示实现推理的迭代优化。Mediom与HIDE共同建立了一个严谨的测试平台和方法论,推动下一代具备文化根基、融入推理提示的多模态习语理解。
原文摘要 · Abstract (English)
Idiomatic reasoning, deeply intertwined with metaphor and culture, remains a blind spot for contemporary language models, whose progress skews toward surface-level lexical and semantic cues. For instance, the Bengali idiom \textit{\foreignlanguage{bengali}{\char"0986\char"0999\char"09CD\char"0997\char"09C1 \char"09B0 \char"09AB\char"09B2 \char"099F\char"0995}} (angur fol tok, ``grapes are sour''): it encodes denial-driven rationalization, yet naive models latch onto the literal fox-and-grape imagery. Addressing this oversight, we present ``Mediom,'' a multilingual, multimodal idiom corpus of 3,533 Hindi, Bengali, and Thai idioms, each paired with gold-standard explanations, cross-lingual translations, and carefully aligned text--image representations. We benchmark both large language models (textual reasoning) and vision-language models (figurative disambiguation) on Mediom, exposing systematic failures in metaphor comprehension. To mitigate these gaps, we propose ``HIDE,'' a Hinting-based Idiom Explanation framework that leverages error-feedback retrieval and targeted diagnostic cues for iterative reasoning refinement. Collectively, Mediom and HIDE establish a rigorous test bed and methodology for culturally grounded, multimodal idiom understanding embedded with reasoning hints in next-generation AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。