arXiv:2511.03908cs.CL2025-11中稿 · NeurIPS

模型在有上下文时能更好理解语言指代,接近人类水平。

Context informs pragmatic interpretation in vision-language models

  • 通过多轮对话测试模型的语境推理能力
  • 有相关上下文时模型表现大幅提升,逼近人类
  • 适合研究人机交互与自然语言理解的学者

迭代指称游戏——参与者反复用语言指认新对象——是检验智能体在多轮语言环境中进行上下文敏感语用推理能力的理想场景。我们测试了人类与视觉-语言模型在不同上下文条件下的表现,包括上下文数量、顺序和相关性。当缺乏相关上下文时,模型表现优于随机水平但显著低于人类;然而,随着相关上下文的引入,模型性能在多轮中显著提升。对于包含抽象指代物的少样本指称游戏,当前机器学习模型仍面临巨大挑战。

原文摘要 · Abstract (English)

Iterated reference games - in which players repeatedly pick out novel referents using language - present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision-language models on trials from iterated reference games, varying the given context in terms of amount, order, and relevance. Without relevant context, models were above chance but substantially worse than humans. However, with relevant context, model performance increased dramatically over trials. Few-shot reference games with abstract referents remain a difficult task for machine learning models.

视觉语言模型语用推理指称游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。