arXiv:2512.09824cs.CVcs.AI2025-12被引 2

通过绑定提示词实现图像视频概念的灵活组合,提升创作一致性与质量。

Composing Concepts from Images and Videos via Concept-prompt Binding

  • 用分层绑定结构将视觉概念映射到提示词,精准拆解复杂概念。
  • 在多个数据集上实现更高的概念一致性与运动质量,优于现有方法。
  • 适合需要跨模态创意生成的研究者和开发者使用。

视觉概念组合旨在将图像和视频中的不同元素融合为一个连贯的视觉输出,但目前仍难以准确提取复杂概念并灵活组合图像与视频中的概念。我们提出 Bind & Compose,一种一次性方法,通过将视觉概念与对应提示词绑定,并从不同来源组合绑定后的提示词,实现灵活的概念组合。该方法采用分层绑定结构,在扩散变换器中通过交叉注意力条件编码视觉概念至提示词,以实现复杂概念的精准分解。为提升概念-提示绑定精度,设计了“多样化吸收机制”,利用额外的吸收令牌消除训练时无关细节的影响。为增强图像与视频概念的兼容性,提出时序解耦策略,通过双分支绑定结构将视频概念训练分为两个阶段进行时序建模。评估表明,本方法在概念一致性、提示保真度和运动质量方面均优于现有方法,为视觉创造力开辟新可能。

原文摘要 · Abstract (English)

Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls short in accurately extracting complex concepts from visual inputs and flexibly combining concepts from both images and videos. We introduce Bind & Compose, a one-shot method that enables flexible visual concept composition by binding visual concepts with corresponding prompt tokens and composing the target prompt with bound tokens from various sources. It adopts a hierarchical binder structure for cross-attention conditioning in Diffusion Transformers to encode visual concepts into corresponding prompt tokens for accurate decomposition of complex visual concepts. To improve concept-token binding accuracy, we design a Diversify-and-Absorb Mechanism that uses an extra absorbent token to eliminate the impact of concept-irrelevant details when training with diversified prompts. To enhance the compatibility between image and video concepts, we present a Temporal Disentanglement Strategy that decouples the training process of video concepts into two stages with a dual-branch binder structure for temporal modeling. Evaluations demonstrate that our method achieves superior concept consistency, prompt fidelity, and motion quality over existing approaches, opening up new possibilities for visual creativity.

概念组合扩散模型跨模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。