提出FlexMUSE框架,实现图文创作中灵活交互与语义对齐。
FlexMUSE: Multimodal Unification and Semantics Enhancement Framework with Flexible interaction for Creative Writing
- 引入可选视觉输入的T2I模块,支持灵活交互模式。
- 通过msaGate和跨模态注意力融合,提升图文语义一致性。
- 适用于需要创意写作与多模态协同的场景。
多模态创意写作(MMCW)旨在生成图文并茂的文章。与常见多模态生成任务如故事生成或图像描述不同,MMCW要求文本与视觉内容之间无严格对应关系,更具抽象性。现有方法强行迁移至该任务时,需特定模态输入或昂贵训练,且常出现模态间语义不一致问题。因此,核心挑战在于以经济方式实现灵活交互下的多模态对齐。本文提出FlexMUSE框架,包含可选视觉输入的T2I模块;通过模态语义对齐门控(msaGate)约束文本输入,促进模态统一;设计基于注意力的跨模态融合以增强特征表示;引入模态语义创意直接偏好优化(mscDPO),通过扩展拒样本提升创作多样性。此外,构建了包含约3000个校准图文对的ArtMUSE数据集。实验表明,FlexMUSE在一致性、创意性和连贯性上均表现优异。
原文摘要 · Abstract (English)
Multi-modal creative writing (MMCW) aims to produce illustrated articles. Unlike common multi-modal generative (MMG) tasks such as storytelling or caption generation, MMCW is an entirely new and more abstract challenge where textual and visual contexts are not strictly related to each other. Existing methods for related tasks can be forcibly migrated to this track, but they require specific modality inputs or costly training, and often suffer from semantic inconsistencies between modalities. Therefore, the main challenge lies in economically performing MMCW with flexible interactive patterns, where the semantics between the modalities of the output are more aligned. In this work, we propose FlexMUSE with a T2I module to enable optional visual input. FlexMUSE promotes creativity and emphasizes the unification between modalities by proposing the modality semantic alignment gating (msaGate) to restrict the textual input. Besides, an attention-based cross-modality fusion is proposed to augment the input features for semantic enhancement. The modality semantic creative direct preference optimization (mscDPO) within FlexMUSE is designed by extending the rejected samples to facilitate the writing creativity. Moreover, to advance the MMCW, we expose a dataset called ArtMUSE which contains with around 3k calibrated text-image pairs. FlexMUSE achieves promising results, demonstrating its consistency, creativity and coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。