arXiv:2509.24410cs.CV2025-09中稿 · WACV 2026 Round 1

5秒生成32张一致多视角图像,提升文本到多视图合成效率

RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis

  • 用时空角度联合隐空间编码视角变化,提升生成一致性
  • 32张多视角图生成仅需约5秒,显著降低延迟
  • 适合需要快速生成3D资产的工业设计与游戏开发场景

从文本提示生成合成多视角图像,是生成合成3D资产的关键桥梁。本文提出RapidMV,一种新型文本到多视图生成模型,可在约5秒内生成32张多视角合成图像。核心在于构建新颖的时空角度隐空间,将空间外观与视角偏移统一编码,以提升效率和多视图一致性。通过分阶段策略化训练流程,实现模型的有效训练。实验表明,RapidMV在一致性与延迟方面优于现有方法,同时保持竞争力的图像质量与图文对齐能力。

原文摘要 · Abstract (English)

Generating synthetic multi-view images from a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic images in just around 5 seconds. In essence, we propose a novel spatio-angular latent space, encoding both the spatial appearance and angular viewpoint deviations into a single latent for improved efficiency and multi-view consistency. We achieve effective training of RapidMV by strategically decomposing our training process into multiple steps. We demonstrate that RapidMV outperforms existing methods in terms of consistency and latency, with competitive quality and text-image alignment.

多视角生成文本生成高效合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。