用视频扩散模型生成多物体合成图,保持物体身份一致且布局自然。
PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
- 基于视频扩散模型,利用时间先验保持物体和背景细节一致。
- 通过合成轨迹数据训练,使物体在推理时自动对齐到目标位置。
- 适合需要高保真多物体合成的广告、影视制作场景。
生成式AI在逼真图像合成方面取得显著进展,但在专业级多物体合成任务中仍存在不足。该任务需同时满足:(i) 各物体身份几乎完全保留,(ii) 背景与色彩精确还原,(iii) 布局与设计元素可控,(iv) 所有物体完整且美观呈现。然而现有顶尖模型常改变物体细节、遗漏或重复物体,导致相对尺寸错误或呈现不一致。为此,我们提出PLACID框架,将一组物体图像转化为视觉上吸引人的多物体合成图。方法上,我们利用带文本控制的预训练图像到视频(I2V)扩散模型,借助视频的时间先验来维持物体一致性、身份及背景细节。其次,提出一种新颖的数据构建策略,生成随机放置物体并平滑移动至目标位置的合成序列,使训练过程与视频模型的时间先验对齐。推理时,物体从随机位置出发,在文本引导下稳定收敛为合理布局,最终帧即为合成图像。大量定量评估与用户研究显示,PLACID在多物体合成任务上优于现有方法,显著提升身份、背景与色彩保真度,减少物体遗漏,结果更美观。
原文摘要 · Abstract (English)
Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of each item's identity, (ii) precise background and color fidelity, (iii) layout and design elements control, and (iv) complete, appealing displays showcasing all objects. However, current state-of-the-art models often alter object details, omit or duplicate objects, and produce layouts with incorrect relative sizing or inconsistent item presentations. To bridge this gap, we introduce PLACID, a framework that transforms a collection of object images into an appealing multi-object composite. Our approach makes two main contributions. First, we leverage a pretrained image-to-video (I2V) diffusion model with text control to preserve objects consistency, identities, and background details by exploiting temporal priors from videos. Second, we propose a novel data curation strategy that generates synthetic sequences where randomly placed objects smoothly move to their target positions. This synthetic data aligns with the video model's temporal priors during training. At inference, objects initialized at random positions consistently converge into coherent layouts guided by text, with the final frame serving as the composite image. Extensive quantitative evaluations and user studies demonstrate that PLACID surpasses state-of-the-art methods in multi-object compositing, achieving superior identity, background, and color preservation, with less omitted objects and visually appealing results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。