arXiv:2510.13702cs.CVcs.AI2025-10被引 3

实现多视角一致的文本定制生成,解决视角与风格统一难题。

MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion

  • 用特征场+时序注意力建模主体几何与身份
  • 多视角生成一致性与定制化效果均优于现有方法
  • 适合需要精准视角控制的视频生成场景

多视角生成中的相机位姿控制与基于提示的定制化是实现可控生成模型的关键。然而,现有模型要么不支持几何一致的定制化,要么缺乏明确的视角控制,难以融合。为此,我们提出新任务——多视角定制化,旨在同时实现多视角位姿控制与个性化定制。由于定制化训练数据稀缺,依赖大规模数据的多视角生成模型难以泛化到多样提示。为此,我们提出基于扩散模型的MVCustom框架,兼顾多视角一致性与定制精度。训练阶段,通过特征场表示学习主体身份与几何结构,并结合增强的文本到视频扩散模型及密集时空注意力,利用时间连贯性保证多视角一致性。推理阶段引入两种新技术:深度感知特征渲染显式强化几何一致性,一致感知潜在补全确保定制主体与背景视角对齐。大量实验表明,MVCustom在多视角一致性与定制化保真度之间达到最佳平衡,有效解决多目标生成难题。

原文摘要 · Abstract (English)

Multi-view generation with camera pose control and prompt-based customization are both essential elements for achieving controllable generative models. However, existing multi-view generation models do not support customization with geometric consistency, whereas customization models lack explicit viewpoint control, making them challenging to unify. Motivated by these gaps, we introduce a novel task, multi-view customization, which aims to jointly achieve multi-view camera pose control and customization. Due to the scarcity of training data in customization, existing multi-view generation models, which inherently rely on large-scale datasets, struggle to generalize to diverse prompts. To address this, we propose MVCustom, a novel diffusion-based framework explicitly designed to achieve both multi-view consistency and customization fidelity. In the training stage, MVCustom learns the subject's identity and geometry using a feature-field representation, incorporating the text-to-video diffusion backbone enhanced with dense spatio-temporal attention, which leverages temporal coherence for multi-view consistency. In the inference stage, we introduce two novel techniques: depth-aware feature rendering explicitly enforces geometric consistency, and consistent-aware latent completion ensures accurate perspective alignment of the customized subject and surrounding backgrounds. Extensive experiments demonstrate that MVCustom achieves the most balanced and consistent competitive performance across multi-view consistency, customization fidelity, demonstrating effective solution of multi-objective generation task.

多视角生成扩散模型定制化几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。