测试大模型能否像人一样用隐含意义沟通,发现它们大多只会直白表达。
Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
- 设计四套新评测,涵盖寓言写作、多智能体游戏等场景
- 顶尖模型在视觉隐喻任务中仍60%使用直白线索
- 能借助共同知识减少直白表达,但难识别未明说的默契
人类交流本质是创造性的,常依赖隐含意义(subtext)——即文字之外的暗示性含义。本文系统研究语言模型是否能在交际场景中使用子文本,并引入四个新评估套件来衡量其能力。评估涵盖寓言创作与理解、受桌游《Dixit》启发的多智能体与多模态游戏。结果表明,前沿模型普遍倾向于过度字面化、显式表达,即便最佳模型在视觉隐喻环境(Visual Allusions)中仍有60%生成直白线索。部分模型可在存在共同知识时减少30%-50%的直白表达,但无法在未明确说明的情况下推断共同知识的存在。寓言理解中,旁注信息和角色设定显著影响子文本解读。本研究为这一高度主观的复杂现象提供了可量化的度量方式,揭示了当前大模型在社交语境下的创造性沟通与推理中的诸多缺陷。希望推动未来面向社会情境的创造性交流研究。
原文摘要 · Abstract (English)
Human communication is fundamentally creative, and often makes use of subtext -- implied meaning that goes beyond the literal content of the text. Here, we systematically study whether language models can use subtext in communicative settings, and introduce four new evaluation suites to assess these capabilities. Our evaluation settings range from writing & interpreting allegories to playing multi-agent and multi-modal games inspired by the rules of board games like Dixit. We find that frontier models generally exhibit a strong bias towards overly literal, explicit communication, and thereby fail to account for nuanced constraints -- even the best performing models generate literal clues 60% of times in one of our environments -- Visual Allusions. However, we find that some models can sometimes make use of common ground with another party to help them communicate with subtext, achieving 30%-50% reduction in overly literal clues; but they struggle at inferring presence of a common ground when not explicitly stated. For allegory understanding, we find paratextual and persona conditions to significantly shift the interpretation of subtext. Overall, our work provides quantifiable measures for an inherently complex and subjective phenomenon like subtext and reveals many weaknesses and idiosyncrasies of current LLMs. We hope this research to inspire future work towards socially grounded creative communication and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。