首个支持文本与布局双控的多对象合成模型,能自动添加互动道具。
Multitwine: Multi-Object Compositing with Text and Layout Control
- 联合训练合成与定制生成,平衡文本与视觉输入
- 可实现从位置关系到复杂动作的多对象交互合成
- 自动生成互动所需辅助物体,适合内容创作与设计
我们提出首个能同时依据文本和布局进行多对象合成的生成模型。该模型可在场景中添加多个对象,捕捉从简单位置关系(如“在...旁边”、“在...前面”)到需重新摆放的复杂动作(如“拥抱”、“弹吉他”)的各类交互。当交互涉及额外道具(如“自拍”)时,模型可自主生成这些支撑对象。通过联合训练合成与主体驱动生成(即定制化),实现了文本与视觉输入的更好融合,显著提升两项任务的性能表现。此外,我们还构建了一条利用视觉与语言模型的自动化数据生成流水线,可高效合成对齐的多模态训练数据。
原文摘要 · Abstract (English)
We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like `taking a selfie', our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。