arXiv:2501.06336cs.CVcs.LG2025-01CVPR被引 98

提出新评估方法,衡量生成图像多视角一致性

MEt3R: Measuring Multi-View Consistency in Generated Images

论文配图:MEt3R: Measuring Multi-View Consistency in Generated Images
图 1 · 摘自论文原文
  • 用DUSt3R实现无监督视角间图像对齐与特征比对
  • 在多个生成模型上验证,发现部分方法存在视角不一致问题
  • 适合研究多视角生成、3D重建的学者参考

我们提出MEt3R,一种用于评估生成图像多视角一致性的度量方法。大规模多视角图像生成模型正在快速推动从稀疏观测中进行3D推断的研究进展。然而,由于生成建模的本质,传统重建指标无法有效衡量生成结果质量,亟需独立于采样过程的评估方法。本文聚焦生成多视角图像间的内在一致性,该评价可脱离具体场景进行。方法采用DUSt3R以前馈方式从图像对中获取稠密3D重建,并将图像内容从一个视角映射到另一个视角。随后比较这些图像的特征图,获得对视角依赖效应不变的相似性分数。利用MEt3R,我们评估了此前大量新型视角与视频生成方法的一致性表现,包括我们开源的多视角隐空间扩散模型。

原文摘要 · Abstract (English)

We introduce MEt3R, a metric for multi-view consistency in generated images. Large-scale generative models for multi-view image generation are rapidly advancing the field of 3D inference from sparse observations. However, due to the nature of generative modeling, traditional reconstruction metrics are not suitable to measure the quality of generated outputs and metrics that are independent of the sampling procedure are desperately needed. In this work, we specifically address the aspect of consistency between generated multi-view images, which can be evaluated independently of the specific scene. Our approach uses DUSt3R to obtain dense 3D reconstructions from image pairs in a feed-forward manner, which are used to warp image contents from one view into the other. Then, feature maps of these images are compared to obtain a similarity score that is invariant to view-dependent effects. Using MEt3R, we evaluate the consistency of a large set of previous methods for novel view and video generation, including our open, multi-view latent diffusion model.

多视角生成图像评估3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。