arXiv:2503.14358cs.CVcs.LG2025-03被引 1

提出RFMI方法,用预训练模型自估计互信息,提升文本图像生成对齐效果。

RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment

  • 基于预训练模型自建互信息估计算法,无需额外数据
  • 通过高互信息样本筛选实现自监督微调,改善图文对齐
  • 无需语言分析或外部评分模型,适合直接部署在现有生成系统

基于流匹配框架训练的修正流(Rectified Flow, RF)模型在文本到图像(T2I)生成任务中已达到顶尖性能。然而多个基准测试显示,生成图像仍存在与提示词对齐不佳的问题,如属性绑定错误、主体位置偏差、数值理解错误等。现有改进方法均针对扩散模型,需依赖辅助数据集、打分模型和提示词语言分析。本文填补此空白:提出RFMI,一种专为RF模型设计的互信息(MI)估计算法,利用预训练模型自身完成估计;进一步提出基于RFMI的自监督微调策略,仅需预训练模型本身即可构建微调数据集——选取预训练模型生成的、图像与提示间点对点互信息高的合成图像。在互信息估计基准上验证了RFMI的有效性,实证表明在SD3.5-Medium上应用该方法可显著提升图文对齐效果,同时保持高质量图像输出。

原文摘要 · Abstract (English)

Rectified Flow (RF) models trained with a Flow matching framework have achieved state-of-the-art performance on Text-to-Image (T2I) conditional generation. Yet, multiple benchmarks show that synthetic images can still suffer from poor alignment with the prompt, i.e., images show wrong attribute binding, subject positioning, numeracy, etc. While the literature offers many methods to improve T2I alignment, they all consider only Diffusion Models, and require auxiliary datasets, scoring models, and linguistic analysis of the prompt. In this paper we aim to address these gaps. First, we introduce RFMI, a novel Mutual Information (MI) estimator for RF models that uses the pre-trained model itself for the MI estimation. Then, we investigate a self-supervised fine-tuning approach for T2I alignment based on RFMI that does not require auxiliary information other than the pre-trained model itself. Specifically, a fine-tuning set is constructed by selecting synthetic images generated from the pre-trained RF model and having high point-wise MI between images and prompts. Our experiments on MI estimation benchmarks demonstrate the validity of RFMI, and empirical fine-tuning on SD3.5-Medium confirms the effectiveness of RFMI for improving T2I alignment while maintaining image quality.

文本生成图像对齐互信息修正流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。