arXiv:2511.12675cs.CV2025-11被引 3

提出新评估框架,让生成视角图更可信。

Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis

  • 用基础模型特征+轻量微调,捕捉视角变换的细微差异。
  • 在三个数据集上验证,分数越低模型越强,与人工偏好一致。
  • 无需参考图也能评估,适合对比不同生成方法。

新视角合成(NVS)旨在从未见视角生成真实图像,但如何判断生成结果是否准确反映预期变换仍是重大挑战。尽管基于扩散模型的生成方法显著提升了质量,现有评估指标仍难以判断生成图像是否既逼真又忠实于源图和视角变换。传统指标如像素级相似度和分布度量常误判错误结果,因未能捕捉源图、视角变化与生成输出间的复杂关系。本文提出一种任务感知评估框架,利用强基线模型Zero123的特征,并通过轻量微调提升判别能力。基于这些特征,引入两个互补指标:基于参考的评分 $D_{\text{PRISM}}$ 与无参考评分 $\text{MMD}_{\text{PRISM}}$。两者均能可靠识别错误生成结果,并与人类偏好研究一致排序模型。在Toys4K、GSO和OmniObject3D三个基准上,对六种NVS方法使用无参考指标,$\text{MMD}_{\text{PRISM}}$ 产生清晰稳定的排名,得分越低代表模型性能越优。

原文摘要 · Abstract (English)

The goal of Novel View Synthesis (NVS) is to generate realistic images of a given content from unseen viewpoints. But how can we trust that a generated image truly reflects the intended transformation? Evaluating its reliability remains a major challenge. While recent generative models, particularly diffusion-based approaches, have significantly improved NVS quality, existing evaluation metrics struggle to assess whether a generated image is both realistic and faithful to the source view and intended viewpoint transformation. Standard metrics, such as pixel-wise similarity and distribution-based measures, often mis-rank incorrect results as they fail to capture the nuanced relationship between the source image, viewpoint change, and generated output. We propose a task-aware evaluation framework that leverages features from a strong NVS foundation model, Zero123, combined with a lightweight tuning step to enhance discrimination. Using these features, we introduce two complementary evaluation metrics: a reference-based score, $D_{\text{PRISM}}$, and a reference-free score, $\text{MMD}_{\text{PRISM}}$. Both reliably identify incorrect generations and rank models in agreement with human preference studies, addressing a fundamental gap in NVS evaluation. Our framework provides a principled and practical approach to assessing synthesis quality, paving the way for more reliable progress in novel view synthesis. To further support this goal, we apply our reference-free metric to six NVS methods across three benchmarks: Toys4K, Google Scanned Objects (GSO), and OmniObject3D, where $\text{MMD}_{\text{PRISM}}$ produces a clear and stable ranking, with lower scores consistently indicating stronger models.

新视角合成评估框架扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。