用视频学习世界动态,统一实现图像生成与编辑。
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
- 将图像任务视为不连续的视频生成,统一处理多种图像任务。
- 从大规模视频中学习真实世界动态,提升光影、姿态等细节一致性。
- 适合需要多任务统一框架的图像生成研究者与开发者。
我们提出UniReal,一个统一框架,用于解决各类图像生成与编辑任务。现有方法通常因任务而异,但本质都遵循一致性和视觉变化的平衡原则。受近期视频生成模型启发,我们提出将图像级任务视为非连续视频生成,将输入输出图像数量视作帧数,从而无缝支持图像生成、编辑、定制、构图等多种任务。尽管面向图像任务,我们利用视频作为可扩展的通用监督来源。UniReal从大规模视频中学习真实世界动态,展现出对阴影、反射、姿态变化和物体交互的高级处理能力,并具备处理新应用的涌现能力。
原文摘要 · Abstract (English)
We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。