arXiv:2607.03752cs.CVcs.AI2026-07

用图像生成反推语言信息,直接评估新兴语言的视觉还原能力

EmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation

论文配图:EmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation
图 1 · 摘自论文原文
  • 用扩散模型从语言消息重建图像,以感知相似度衡量视觉信息保留程度
  • 在MS-COCO数据集上验证,比现有四种方法更准确捕捉视觉内容
  • 适合研究语言演化、多智能体通信或视觉-语言对齐的学者使用

衡量新兴语言在多大程度上编码输入图像的视觉内容仍是一个开放问题。我们称之为‘视觉反射’:即语言消息在不依赖说话者-听者配对的情况下,能否保留可恢复的原始图像信息。现有指标仅通过人类定义的概念词表、自然语言描述、结构距离相关性或参照游戏准确率等间接代理来评估,这些方法可能遗漏真实包含的视觉信息,或错误地赋予无关内容。为此,我们提出EmCom-Diffusion,一种直接评估视觉反射的框架:它利用预训练文本到图像扩散模型,基于(图像,新兴语言消息)配对进行微调,通过重建图像与原图之间的感知相似度评分来衡量视觉反射。我们在MS-COCO数据集上,结合参照游戏任务,对三种预训练视觉编码器下的随机基线和固定词基线进行了验证,并与四种现有指标(CBM、监督翻译、TopSim、R@1)进行了对比。结果表明,EmCom-Diffusion能捕捉其他方法遗漏的视觉内容。

原文摘要 · Abstract (English)

Measuring the extent to which emergent languages encode the visual content of their inputs is an open problem. We refer to this property as visual reflection: the extent to which emergent messages preserve information about their source images that can be recovered without appeal to the speaker-listener pair that produced them. Existing metrics measure it only indirectly, through proxies such as human-defined concept inventories, natural-language captions, structural distance correlations, or Referential Game accuracy, each of which can either miss visual content the message encodes or credit content it does not. We propose EmCom-Diffusion, an evaluation framework that measures visual reflection directly: it reconstructs each input image from its emergent message and compares the reconstruction with the original image itself, rather than with human-defined targets. Concretely, it finetunes a pretrained text-to-image diffusion model on (image, emergent-message) pairs and scores visual reflection as the perceptual similarity between the reconstructed and original images, operating generatively rather than discriminatively. Instantiating it on MS-COCO with a Referential Game, we validate the metric against random and fixed-token baselines under three pretrained visual encoders, and compare it against four existing metrics (CBM, supervised translation, TopSim, and R@1). EmCom-Diffusion captures visual content the other metrics miss.

语言演化视觉反射扩散模型多智能体通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。