arXiv:2507.00006cs.GRcs.LG2025-07ICCV被引 8

构建多视角生成评估基准,突破传统对比真值的局限。

MVGBench: Comprehensive Benchmark for Multi-view Generation Models

  • 用3D自一致性度量评估多视角生成的几何纹理一致性。
  • 12种模型在4个数据集上测试,发现泛化与鲁棒性是主要短板。
  • 提出ViFiGen新方法,在3D一致性上超越所有现有模型。

我们提出MVGBench,一个全面评估多视角图像生成模型(MVGs)的基准,涵盖三维一致性(几何与纹理)、图像质量及语义准确性(使用视觉语言模型)。近期,MVGs已成为3D物体生成的主要驱动力。然而,现有度量方法将生成图像与真实目标视图对比,不适用于生成任务中存在多种合理解的情况。此外,不同MVGs在不同视角、合成数据和光照条件下训练,其对这些因素的鲁棒性及向真实数据的泛化能力极少被充分评估。缺乏严格评估协议,也难以判断哪些设计选择推动了进展。MVGBench从三个维度评估:最优设置性能、真实数据泛化能力与鲁棒性。我们引入新颖的3D自一致性度量,通过不重叠生成多视角重建进行比较,而非依赖真值。系统性地在4个精心筛选的真实与合成数据集上对比12种现有MVGs。分析揭示现有方法在鲁棒性和泛化方面的重要局限,并识别出最关键的若干设计选择。基于发现的最佳实践,我们提出ViFiGen,该方法在3D一致性上优于所有评估过的MVGs。代码、模型与基准套件将公开发布。

原文摘要 · Abstract (English)

We propose MVGBench, a comprehensive benchmark for multi-view image generation models (MVGs) that evaluates 3D consistency in geometry and texture, image quality, and semantics (using vision language models). Recently, MVGs have been the main driving force in 3D object creation. However, existing metrics compare generated images against ground truth target views, which is not suitable for generative tasks where multiple solutions exist while differing from ground truth. Furthermore, different MVGs are trained on different view angles, synthetic data and specific lightings -- robustness to these factors and generalization to real data are rarely evaluated thoroughly. Without a rigorous evaluation protocol, it is also unclear what design choices contribute to the progress of MVGs. MVGBench evaluates three different aspects: best setup performance, generalization to real data and robustness. Instead of comparing against ground truth, we introduce a novel 3D self-consistency metric which compares 3D reconstructions from disjoint generated multi-views. We systematically compare 12 existing MVGs on 4 different curated real and synthetic datasets. With our analysis, we identify important limitations of existing methods specially in terms of robustness and generalization, and we find the most critical design choices. Using the discovered best practices, we propose ViFiGen, a method that outperforms all evaluated MVGs on 3D consistency. Our code, model, and benchmark suite will be publicly released.

多视角生成3D一致性评估基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。