研究扩散修复缺陷如何影响图文模型的文本生成质量。
How Do Inpainting Artifacts Propagate to Language?
- 分两阶段诊断:先修复图像再生成描述,控制变量对比效果。
- 修复精度越高,文字描述在词汇和语义上越准确。
- 发现修复错误会引发模型各层行为系统性偏差,适合多模态研究者参考。
我们研究基于扩散模型的图像修复引入的视觉瑕疵如何影响视觉语言模型的语言生成能力。采用两阶段诊断框架:先对图像掩码区域进行重建,再输入至描述生成模型,从而在多个数据集上可控地比较原始与重建输入生成的描述。分析显示,像素级与感知级重建指标与词汇和语义层面的描述质量存在一致关联。进一步分析中间视觉表征与注意力模式发现,修复瑕疵会导致模型行为出现系统性、层级依赖性变化。这些结果共同构建了一个实用的诊断框架,用于评估视觉重建质量如何影响多模态系统中的语言生成。
原文摘要 · Abstract (English)
We study how visual artifacts introduced by diffusion-based inpainting affect language generation in vision-language models. We use a two-stage diagnostic setup in which masked image regions are reconstructed and then provided to captioning models, enabling controlled comparisons between captions generated from original and reconstructed inputs. Across multiple datasets, we analyze the relationship between reconstruction fidelity and downstream caption quality. We observe consistent associations between pixel-level and perceptual reconstruction metrics and both lexical and semantic captioning performance. Additional analysis of intermediate visual representations and attention patterns shows that inpainting artifacts lead to systematic, layer-dependent changes in model behavior. Together, these results provide a practical diagnostic framework for examining how visual reconstruction quality influences language generation in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。