arXiv:2607.00601cs.CL2026-07

测试大模型在禁词约束下的描述能力,发现其表现远不如人类。

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo

论文配图:"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo
图 1 · 摘自论文原文
  • 通过提示、生成约束到内部表征干预,层层测试模型表现。
  • 模型在遵守规则和有效传达概念间存在权衡,且远逊于人类猜词能力。
  • 揭示了当前模型在约束条件下语义理解的不足,适合研究语言生成与约束的学者参考。

Taboo游戏要求描述目标词时避开一组禁用词,使其他玩家能猜出它。这一看似简单的任务融合了严格的词汇约束与有效的沟通需求,是检验大语言模型在推理时应对多重挑战的理想场景。我们评估了两种开源模型,在从提示到生成时约束再到内部表征干预的不同层级条件下,对输出进行检测:禁用词违规率、以大模型为裁判的猜测成功率(针对人类与机器猜题者),以及模型策略是否符合人类玩家行为。结果显示,不同条件下的规则遵守与沟通有效性存在不同权衡,且模型作为猜题者的性能显著低于人类,表明在约束下实现词汇语义接地仍是当前大模型的开放挑战。

原文摘要 · Abstract (English)

The game of Taboo requires describing a target word without using a set of forbidden words, so that other players can guess it. This deceptively simple task combines strict lexical constraints with the need for communicatively effective descriptions, making it a compelling playground for examining how LLMs navigate competing demands at inference time. We evaluate two open-weight models under conditions that intervene at progressively deeper levels of the generative process, from prompting to generation-time constraints to internal representations manipulations. We assess their outputs through forbidden word violation detection, LLM-as-a-judge measuring the degree to which generated descriptions successfully evoke the target concept for both human and machine guessers, and examining whether the strategies models adopt under constraint align with those of human players. Our results show that compliance with the rules of the game and communicative effectiveness trade off differently across conditions, and that models remain substantially weaker than humans as guessers, suggesting that lexical grounding under constraint is an open challenge for current language models.

语言模型约束生成人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。