通过语义批评与频域对齐,提升文本生成图像的准确性和细节。
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
- 用多模型分析提示语,生成高层语义反馈指导生成。
- 在频域融合中间结果,保留细节同时增强结构一致性。
- 无需训练,可插件式接入现有生成模型,适合追求高保真图像的用户。
近期文本到图像扩散模型虽具备出色视觉保真度,但在复杂提示下的语义对齐能力仍不足。本文提出 CritiFusion,一种推理阶段的框架,结合多模态语义批评机制与频域精修,提升文本-图像一致性与细节表现。CritiCore 模块利用视觉语言模型和多个大语言模型丰富提示上下文,生成高层语义反馈,引导扩散过程更贴合提示意图。SpecFusion 在频域融合中间生成状态,注入粗粒度结构信息的同时保留高频细节。该方法无需额外训练,可作为插件式优化模块兼容现有扩散模型。标准基准测试表明,该方法显著提升人机评估中的文本对应度与视觉质量,人类偏好评分和审美评价均优于多数基线,达到与先进奖励优化方法相当的水平。定性结果进一步显示其在细节、真实感和提示忠实度上的优势,验证了语义批评与频域对齐策略的有效性。
原文摘要 · Abstract (English)
Recent text-to-image diffusion models have achieved remarkable visual fidelity but often struggle with semantic alignment to complex prompts. We introduce CritiFusion, a novel inference-time framework that integrates a multimodal semantic critique mechanism with frequency-domain refinement to improve text-to-image consistency and detail. The proposed CritiCore module leverages a vision-language model and multiple large language models to enrich the prompt context and produce high-level semantic feedback, guiding the diffusion process to better align generated content with the prompt's intent. Additionally, SpecFusion merges intermediate generation states in the spectral domain, injecting coarse structural information while preserving high-frequency details. No additional model training is required. CritiFusion serves as a plug-in refinement stage compatible with existing diffusion backbones. Experiments on standard benchmarks show that our method notably improves human-aligned metrics of text-to-image correspondence and visual quality. CritiFusion consistently boosts performance on human preference scores and aesthetic evaluations, achieving results on par with state-of-the-art reward optimization approaches. Qualitative results further demonstrate superior detail, realism, and prompt fidelity, indicating the effectiveness of our semantic critique and spectral alignment strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。