用流模型实现跨模态样本结构对齐,速度快质量高。
ATATA: One Algorithm to Align Them All
- 基于修正流模型联合运输样本片段,实现结构对齐
- 图像与视频生成超越现有水平,3D生成快百倍且质量相当
- 适用于图像、视频、3D生成,适合追求高效对齐的开发者
我们提出一种新的多模态算法,用于联合推断具有结构对齐特性的配对样本,采用修正流模型。现有方法虽有耦合生成过程,但未从结构对齐视角出发。近期工作使用得分蒸馏采样生成对齐3D模型,但该方法耗时长、易模式崩溃,常得卡通化结果。相比之下,我们的方法依赖样本空间中一段的联合传输,推理速度更快。该方法可构建于任意在结构化潜在空间运行的修正流模型之上。我们将其应用于图像、视频和3D形状生成领域,采用最先进基线进行评估,并与基于编辑和联合推断的方法对比。实验表明,本方法生成的样本对具有高度结构对齐性,且视觉质量优异。在图像与视频生成任务中,性能超越现有最佳水平;在3D生成中,质量相当,但速度提升数个数量级。
原文摘要 · Abstract (English)
We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing methods propose a codependent generation process, they do not view the problem of joint generation from a structural alignment perspective. Recent work uses Score Distillation Sampling to generate aligned 3D models, but SDS is known to be time-consuming, prone to mode collapse, and often provides cartoonish results. By contrast, our suggested approach relies on the joint transport of a segment in the sample space, yielding faster computation at inference time. Our approach can be built on top of an arbitrary Rectified Flow model operating on the structured latent space. We show the applicability of our method to the domains of image, video, and 3D shape generation using state-of-the-art baselines and evaluate it against both editing-based and joint inference-based competing approaches. We demonstrate a high degree of structural alignment for the sample pairs obtained with our method and a high visual quality of the samples. Our method improves the state-of-the-art for image and video generation pipelines. For 3D generation, it is able to show comparable quality while working orders of magnitude faster.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。