用跨模态注意力蒸馏,让新视角图像与几何同步生成。
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
- 通过变形与修补框架,将新视角生成转为图像和几何的联合修复任务。
- 在未见场景上实现高质量图像与几何外推,生成结果对齐度高。
- 适合做3D重建、新视角合成的研究者,尤其关注几何一致性问题。
我们提出一种基于扩散模型的框架,通过变形与修补方法实现新视角图像与几何的对齐生成。不同于需要密集姿态图像或仅限域内视角的已有方法,本方法利用现成的几何预测器,从参考图像中预测部分几何信息,并将新视角合成建模为图像与几何的联合修复任务。为确保生成图像与几何的精准对齐,我们提出跨模态注意力蒸馏机制,在训练和推理阶段将图像扩散分支的注意力图注入并行的几何扩散分支。该多任务策略产生协同效应,提升几何鲁棒性图像生成能力及清晰几何预测。进一步引入基于邻近性的网格条件机制,融合深度与法向线索,介于点云之间并过滤错误预测的几何干扰生成过程。实验表明,该方法在多种未见场景中均实现高保真外推式视图合成,插值设置下具备竞争力的重建质量,并生成几何对齐的彩色点云,支持完整3D补全。项目主页见 https://cvlab-kaist.github.io/MoAI。
原文摘要 · Abstract (English)
We introduce a diffusion-based framework that performs aligned novel view image and geometry generation via a warping-and-inpainting methodology. Unlike prior methods that require dense posed images or pose-embedded generative models limited to in-domain views, our method leverages off-the-shelf geometry predictors to predict partial geometries viewed from reference images, and formulates novel-view synthesis as an inpainting task for both image and geometry. To ensure accurate alignment between generated images and geometry, we propose cross-modal attention distillation, where attention maps from the image diffusion branch are injected into a parallel geometry diffusion branch during both training and inference. This multi-task approach achieves synergistic effects, facilitating geometrically robust image synthesis as well as well-defined geometry prediction. We further introduce proximity-based mesh conditioning to integrate depth and normal cues, interpolating between point cloud and filtering erroneously predicted geometry from influencing the generation process. Empirically, our method achieves high-fidelity extrapolative view synthesis on both image and geometry across a range of unseen scenes, delivers competitive reconstruction quality under interpolation settings, and produces geometrically aligned colored point clouds for comprehensive 3D completion. Project page is available at https://cvlab-kaist.github.io/MoAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。