评测多模态语境下模型理解习语的能力,推动跨语言习语表征研究。
SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation

- 设计双子任务:图像匹配与序列预测,评估模型对习语的多模态理解。
- 顶尖方法融合预训练大模型,在专家混合架构中实现人类级表现。
- 适合关注多模态语义、跨语言理解的研究者和开发者。
习语在自然语言处理中构成独特挑战,因其含义无法从字面词语直接推断。尽管大型语言模型(LLMs)取得进展,习语理解仍是构建鲁棒语义表征的重大障碍。我们为 SemEval-2025 第1任务:AdMiRe(Advancing Multimodal Idiomaticity Representation)提出数据集与任务,旨在评估并提升模型在多模态场景和多语言环境下对习语的理解能力。参赛者需完成两项子任务:根据图像与习语或字面意义的匹配程度进行排序,以及预测图像序列中的下一帧。最优方法通过在专家混合设置中结合预训练语言与视觉-语言模型,并使用多个查询来弥补模型在习语表征上的不足,最终达到人类水平性能。
原文摘要 · Abstract (English)
Idiomatic expressions present a unique challenge in NLP, as their meanings are often not directly inferable from their constituent words. Despite recent advancements in Large Language Models (LLMs), idiomaticity remains a significant obstacle to robust semantic representation. We present datasets and tasks for SemEval-2025 Task 1: AdMiRe (Advancing Multimodal Idiomaticity Representation), which challenges the community to assess and improve models' ability to interpret idiomatic expressions in multimodal contexts and in multiple languages. Participants competed in two subtasks: ranking images based on their alignment with idiomatic or literal meanings, and predicting the next image in a sequence. The most effective methods achieved human-level performance by leveraging pretrained LLMs and vision-language models in mixture-of-experts settings, with multiple queries used to smooth over the weaknesses in these models' representations of idiomaticity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。