研究文字描述如何破坏图像生成质量,发现99%的图像严重失真。
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
- 用文字转述图像再还原,测试信息丢失程度
- 99.3%图像感知质量下降,91.5%结构信息丢失
- 适合关注多模态系统缺陷的研究者
随着多模态AI在创意流程中的广泛应用,理解视觉-语言-视觉链路中的信息损失变得尤为重要。然而,视觉内容经由文本中间表示时的退化程度仍缺乏量化。本文对描述-生成瓶颈进行了实证分析,通过该管道生成了150组图像对,并使用LPIPS、SSIM和颜色距离等指标,在感知、结构和色彩维度上评估信息保留情况。结果表明,99.3%的样本出现显著的感知退化,91.5%显示明显的结构信息丢失,为当前多模态系统中描述-生成瓶颈的存在提供了可量化的实证依据。
原文摘要 · Abstract (English)
With the increasing integration of multimodal AI systems in creative workflows, understanding information loss in vision-language-vision pipelines has become important for evaluating system limitations. However, the degradation that occurs when visual content passes through textual intermediation remains poorly quantified. In this work, we provide empirical analysis of the describe-then-generate bottleneck, where natural language serves as an intermediate representation for visual information. We generated 150 image pairs through the describe-then-generate pipeline and applied existing metrics (LPIPS, SSIM, and color distance) to measure information preservation across perceptual, structural, and chromatic dimensions. Our evaluation reveals that 99.3% of samples exhibit substantial perceptual degradation and 91.5% demonstrate significant structural information loss, providing empirical evidence that the describe-then-generate bottleneck represents a measurable and consistent limitation in contemporary multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。