让文字描述精准控制图像风格转换,保持结构不变。
Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance
- 用语言提示引导每类属性的视觉变化,实现细粒度控制。
- 在CelebA(Dialog)和BDD100K上实现高保真与结构一致性。
- 适合需要可控、可解释图像生成的跨模态应用者。
多领域图像到图像转换需将自然语言提示中的语义差异映射为对应视觉变换,同时保留无关结构与语义内容。现有方法难以维持结构完整性,且在多领域场景下缺乏细粒度属性级控制。本文提出LACE(语言接地属性可控翻译),基于两个组件:(1) GLIP-Adapter融合全局语义与局部结构特征以保持一致性;(2) 多领域控制引导机制,显式将源与目标提示间的语义差值转化为各属性的翻译向量,使语言语义与域级视觉变化对齐。二者协同实现属性独立强度调节的组合式多域控制。在CelebA(Dialog)与BDD100K上的实验表明,LACE在视觉保真度、结构保留与可解释的领域特异性控制方面均优于先前基线,定位为连接语言语义与可控视觉转换的跨模态生成框架。
原文摘要 · Abstract (English)
Multi-domain image-to-image translation re quires grounding semantic differences ex pressed in natural language prompts into corresponding visual transformations, while preserving unrelated structural and seman tic content. Existing methods struggle to maintain structural integrity and provide fine grained, attribute-specific control, especially when multiple domains are involved. We propose LACE (Language-grounded Attribute Controllable Translation), built on two compo nents: (1) a GLIP-Adapter that fuses global semantics with local structural features to pre serve consistency, and (2) a Multi-Domain Control Guidance mechanism that explicitly grounds the semantic delta between source and target prompts into per-attribute translation vec tors, aligning linguistic semantics with domain level visual changes. Together, these modules enable compositional multi-domain control with independent strength modulation for each attribute. Experiments on CelebA(Dialog) and BDD100K demonstrate that LACE achieves high visual fidelity, structural preservation, and interpretable domain-specific control, surpass ing prior baselines. This positions LACE as a cross-modal content generation framework bridging language semantics and controllable visual translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。