arXiv:2502.05165cs.CV2025-02CVPR被引 10

首个支持文本与布局双控的多对象合成模型,能自动添加互动道具。

Multitwine: Multi-Object Compositing with Text and Layout Control

  • 联合训练合成与定制生成,平衡文本与视觉输入
  • 可实现从位置关系到复杂动作的多对象交互合成
  • 自动生成互动所需辅助物体,适合内容创作与设计

我们提出首个能同时依据文本和布局进行多对象合成的生成模型。该模型可在场景中添加多个对象,捕捉从简单位置关系(如“在...旁边”、“在...前面”)到需重新摆放的复杂动作(如“拥抱”、“弹吉他”)的各类交互。当交互涉及额外道具(如“自拍”)时,模型可自主生成这些支撑对象。通过联合训练合成与主体驱动生成(即定制化),实现了文本与视觉输入的更好融合,显著提升两项任务的性能表现。此外,我们还构建了一条利用视觉与语言模型的自动化数据生成流水线,可高效合成对齐的多模态训练数据。

原文摘要 · Abstract (English)

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like `taking a selfie', our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data.

图像合成多对象生成文本控制布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。