arXiv:2511.21698cs.MMcs.AI2025-11

通过文本图像原型建模,提升多模态生成的语义一致性和风格统一性。

TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement

  • 基于文本、图像和视觉原型信号,设计双对齐注意力与差异操作模块
  • 在多个评估指标上超越基线,尤其在创意性和语义一致性上显著提升
  • 适合需要高质量多模态内容生成的研究与应用,如艺术创作、广告设计

多模态生成难以保证主题连贯性与风格一致性。现有方法存在跨模态不匹配问题,缺乏对共性与差异的显式建模,依赖细粒度训练的方法又难以兼顾语义精度与写作风格一致性,导致生成质量不佳。为此,我们提出 extbf{ extit{TIPPo}},一个具备显式输入建模与综合优化目标的简洁高效框架。该框架通过多模态编码器与适配器提取输入文本与图像,并计算视觉原型。随后,文本、图像与原型信号输入至所提出的双对齐注意力与差异操作模块,再经语言模型解码。其中, extbf{Po}lishPPO 增强风格一致性,无监督对比学习在 SFT 阶段缓解样本间表示坍缩。实验结果表明, extbf{ extit{TIPPo}} 在自动评估与基于大语言模型的创意性及语义一致性判别中均表现优异。

原文摘要 · Abstract (English)

Multi-modal generation struggles to ensure thematic coherence and style consistency. Semantically, existing methods suffer from cross-modal mismatch and lack explicit modeling of commonality and discrepancy. Methods that rely on fine-grained training fail to balance semantic precision with writing style consistency. These shortcomings lead to suboptimal generation quality. To tackle these issues, we propose \textbf{\textit{TIPPo}}, a simple yet effective framework with explicit input modeling and comprehensive optimization objectives. It extracts the input text and images via multi-modal encoder and adapters, then measures the visual prototype. \textbf{T}extual, \textbf{I}mage, and \textbf{P}rototype signals are then fed to our proposed Dual Alignment Attention and Difference Operator modules before language model decoding. The proposed \textbf{Po}lishPPO reinforces the style consistency, while the unsupervised contrastive learning during SFT mitigates inter-sample representation collapse. Experimental results demonstrate the promising performance of \textbf{\textit{TIPPo}} in automatic evaluation and LLM-based criteria for creativity and semantic consistency.

多模态生成风格一致性原型建模LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。