通过多轮对话让模型逐步修正图像,自动评估重建效果。
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

- 用多轮对话迭代优化图像生成,每轮输出可直观观察改进过程。
- 描述器决定重建质量,生成器影响是否能通过迭代提升性能。
- 数学与几何图像最难重建,强描述器会使用更丰富的修正词汇。
我们提出图像重建游戏,一个全自动基准测试,视觉-语言模型通过多轮交互向图像生成器发出修正指令,使共同理解过程以生成图像的形式直接可见。在七个图像类别上对两种描述器模型与两种生成器模型进行组合测试,发现描述器是重建质量的主要决定因素,而生成器决定了迭代优化是否有效。数学与几何类图像最具挑战性。描述器的词汇预算显著影响收敛:短预算导致初始图像稀疏,留有较大改进空间;长预算虽提升绝对质量,但剩余可修正内容较少。强描述器使用涵盖空间、数值和结构类的丰富修正词汇,弱描述器则聚焦表面属性,常在几轮后停止。人工验证显示,最佳自动化评判器与人类偏好仅达到轻微至中等一致,自动化评分需经人工校准才可可靠使用。
原文摘要 · Abstract (English)
We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns, making accumulated common ground directly observable as a rendered image. Benchmarking two Describer models crossed with two Generator models across seven image categories, we find that the describer is the dominant factor in reconstruction quality, while the generator determines whether iterative refinement helps or hurts. Mathematical and geometric images pose the greatest challenge. The describer's token budget strongly affects convergence: shorter budgets yield sparser first renderings with more room for visible improvement, while longer budgets raise absolute quality but leave less to fix. Stronger describers use a richer correction vocabulary spanning spatial, numeric, and structural categories, while weaker describers concentrate on surface properties and tend to stop after a few turns. Human validation shows that the best automated judge reaches only slight-to-fair agreement with human preferences, and automated scores require human recalibration to be used reliably.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。