发现图文模型间风格差异,揭示文本生成更易识别来源。
Asymmetric Idiosyncrasies in Multimodal Models
- 用分类法检测图文模型的风格指纹,对比跨模态差异。
- 文本分类准确率达99.70%,图像分类最高仅50%。
- 适合关注多模态对齐与提示遵循能力的研究者。
本文研究了图文生成模型中的风格差异及其对文本到图像模型的下游影响。设计了一种系统性分析方法:给定生成的文本或对应图像,训练神经网络预测其来源的文本生成模型。结果显示,文本分类准确率高达99.70%,表明文本模型嵌入了显著的风格特征。相比之下,这些特征在生成图像中几乎消失,即使采用最先进的Flux模型,分类准确率也降至最多50%。进一步分析发现,生成图像未能保留原文本中的关键差异,如细节层次、颜色纹理侧重以及场景中物体分布等。整体而言,该基于分类的框架为量化文本模型的风格特征和文本到图像系统的提示遵循能力提供了新方法。
原文摘要 · Abstract (English)
In this work, we study idiosyncrasies in the caption models and their downstream impact on text-to-image models. We design a systematic analysis: given either a generated caption or the corresponding image, we train neural networks to predict the originating caption model. Our results show that text classification yields very high accuracy (99.70\%), indicating that captioning models embed distinctive stylistic signatures. In contrast, these signatures largely disappear in the generated images, with classification accuracy dropping to at most 50\% even for the state-of-the-art Flux model. To better understand this cross-modal discrepancy, we further analyze the data and find that the generated images fail to preserve key variations present in captions, such as differences in the level of detail, emphasis on color and texture, and the distribution of objects within a scene. Overall, our classification-based framework provides a novel methodology for quantifying both the stylistic idiosyncrasies of caption models and the prompt-following ability of text-to-image systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。