让AI理解并迁移视觉隐喻的抽象逻辑,生成更深层创意内容。
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
- 用认知理论构建结构化隐喻框架,分离抽象逻辑与具体图像。
- 多智能体协作实现隐喻解构、迁移与高质量生成,准确率超基线。
- 适合广告、媒体等需要高阶创意的场景,可自动生成有深度视觉表达的内容。
视觉隐喻是一种高层次的人类创造力形式,通过跨域语义融合将抽象概念转化为有力的视觉修辞。尽管生成式AI取得显著进展,现有模型仍局限于像素级指令对齐和表面外观保持,无法捕捉真实隐喻生成所需的底层抽象逻辑。为此,我们提出视觉隐喻迁移(VMT)任务,要求模型自主剥离参考图像中的“创意本质”,并将该抽象逻辑重新投射到用户指定的目标主体上。我们设计了一种受认知启发的多智能体框架,通过新型模式语法(G)实现概念融合理论(CBT),以结构化方式解耦关系不变性与具体视觉实体,为跨域逻辑重构提供严谨基础。整个流程由感知代理提取参考图的模式,转移代理维持通用空间不变性以发现合适载体,生成代理完成高保真合成,诊断代理则模拟专业评审,通过闭环回溯识别并修正抽象逻辑、组件选择与提示编码中的错误。大量实验与人工评估表明,本方法在隐喻一致性、类比恰当性和视觉创造性方面显著优于现有最先进基线,为广告与媒体领域的自动化高影响力创意应用开辟新路径。项目页面含源码与自包含技能:https://yuci-gpt.github.io/Beyond-Pixels/。
原文摘要 · Abstract (English)
A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite the remarkable progress of generative AI, existing models remain largely confined to pixel-level instruction alignment and surface-level appearance preservation, failing to capture the underlying abstract logic necessary for genuine metaphorical generation. To bridge this gap, we introduce the task of Visual Metaphor Transfer (VMT), which challenges models to autonomously decouple the "creative essence" from a reference image and re-materialize that abstract logic onto a user-specified target subject. We propose a cognitive-inspired, multi-agent framework that operationalizes Conceptual Blending Theory (CBT) through a novel Schema Grammar ("G"). This structured representation decouples relational invariants from specific visual entities, providing a rigorous foundation for cross-domain logic re-instantiation. Our pipeline executes VMT through a collaborative system of specialized agents: a perception agent that distills the reference into a schema, a transfer agent that maintains generic space invariance to discover apt carriers, a generation agent for high-fidelity synthesis and a hierarchical diagnostic agent that mimics a professional critic, performing closed-loop backtracking to identify and rectify errors across abstract logic, component selection, and prompt encoding. Extensive experiments and human evaluations demonstrate that our method significantly outperforms SOTA baselines in metaphor consistency, analogy appropriateness, and visual creativity, paving the way for automated high-impact creative applications in advertising and media. Project page with source code and self-contained skills is at https://yuci-gpt.github.io/Beyond-Pixels/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。