对比显式与隐式提示,发现AI难从隐式提示中理解沟通效率
Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

- 显式提示下模型能生成高效指代表达
- 隐式提示下相同模型无法自发追求沟通效率
- 揭示人与AI在沟通意图理解上的本质差异
两项近期研究(Jones等,2026;Zeng等,2026)对视觉语言大模型(LVLMs)能否协调生成高效指代表达得出看似矛盾的结论。本文在控制任务差异的前提下,直接比较两种提示风格。结果复现了显式提示下模型可实现高效指代表达的发现,表明先前结果差异并非由任务设计引起。但同时发现,相同模型在隐式提示下无法推断出需要追求沟通效率,凸显人类与人工智能在沟通意图理解上的关键差异。
原文摘要 · Abstract (English)
Two recent studies (Jones et al. (2026); Zeng et al. (2026)) reach apparently contradictory conclusions about whether LVLMs can coordinate on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when explicitly prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more implicit prompt, highlighting critical differences between how humans and AI systems communicate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。