arXiv:2410.18823cs.CVcs.AI2024-10NeurIPS被引 1

跨语言视觉文字设计迁移,让海报字体风格精准转换

Towards Visual Text Design Transfer Across Languages

  • 不依赖风格描述,用符号潜空间与OCR反馈优化多语言字体生成
  • 在跨语言字体迁移任务中,显著提升风格一致性和可读性
  • 适合做多语言视觉内容生成的研究者和设计师参考

视觉文字设计在电影海报、专辑封面等多模态场景中对主题、情感和氛围传达至关重要。跨语言迁移不仅涉及文本翻译,还需适配美学与风格特征。为此,我们提出多模态风格迁移任务(MuST-Bench),用于评估视觉文字生成模型在不同书写系统间保持设计意图的能力。初步实验表明,现有模型因文本描述无法充分表达视觉风格而表现不佳。为此,我们提出SIGIL框架,通过三个创新实现无需风格描述的多模态风格迁移:多语言字符潜空间、预训练VAE提供稳定风格引导,以及结合强化学习反馈的OCR模型优化可读字符生成。SIGIL在风格一致性、可读性和视觉保真度上均优于基线模型。我们已公开MuST-Bench数据集供研究使用(https://huggingface.co/datasets/yejinc/MuST-Bench)。

原文摘要 · Abstract (English)

Visual text design plays a critical role in conveying themes, emotions, and atmospheres in multimodal formats such as film posters and album covers. Translating these visual and textual elements across languages extends the concept of translation beyond mere text, requiring the adaptation of aesthetic and stylistic features. To address this, we introduce a novel task of Multimodal Style Translation (MuST-Bench), a benchmark designed to evaluate the ability of visual text generation models to perform translation across different writing systems while preserving design intent. Our initial experiments on MuST-Bench reveal that existing visual text generation models struggle with the proposed task due to the inadequacy of textual descriptions in conveying visual design. In response, we introduce SIGIL, a framework for multimodal style translation that eliminates the need for style descriptions. SIGIL enhances image generation models through three innovations: glyph latent for multilingual settings, pretrained VAEs for stable style guidance, and an OCR model with reinforcement learning feedback for optimizing readable character generation. SIGIL outperforms existing baselines by achieving superior style consistency and legibility while maintaining visual fidelity, setting itself apart from traditional description-based approaches. We release MuST-Bench publicly for broader use and exploration https://huggingface.co/datasets/yejinc/MuST-Bench.

多模态生成跨语言视觉文字风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。