arXiv:2511.07193cs.CL2025-11AAAI

测试大模型在微妙语境中理解表情符号歧义的能力

EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language Models

  • 构建含对比语境的 emoji 歧义测试集
  • 强模型在细微语境下仍常判断错误
  • 揭示模型对语用对比不敏感的缺陷

大型语言模型(LLMs)日益应用于实际通信场景,但其处理上下文依赖歧义的能力仍缺乏深入研究。本文提出 EMODIS,一个评估 LLM 在极简但具有对比性的文本语境中解析模糊表情符号表达能力的新基准。EMODIS 中每条数据包含一个含表情符号的歧义句子、两个导致不同解释的对立语境,以及需语境推理的问题。我们评估了开源与 API 型 LLM,发现即使最强模型在仅有细微语境线索时也频繁无法区分含义。进一步分析显示,模型存在对主流解释的系统性偏见,且对语用对比敏感度有限。EMODIS 提供了一个严格的上下文消歧测试平台,凸显了人类与 LLM 在语义推理上的差距。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in real-world communication settings, yet their ability to resolve context-dependent ambiguity remains underexplored. In this work, we present EMODIS, a new benchmark for evaluating LLMs' capacity to interpret ambiguous emoji expressions under minimal but contrastive textual contexts. Each instance in EMODIS comprises an ambiguous sentence containing an emoji, two distinct disambiguating contexts that lead to divergent interpretations, and a specific question that requires contextual reasoning. We evaluate both open-source and API-based LLMs, and find that even the strongest models frequently fail to distinguish meanings when only subtle contextual cues are present. Further analysis reveals systematic biases toward dominant interpretations and limited sensitivity to pragmatic contrast. EMODIS provides a rigorous testbed for assessing contextual disambiguation, and highlights the gap in semantic reasoning between humans and LLMs.

表情符号语义消歧大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。