arXiv:2501.06488cs.CVcs.AI2025-01TPAMI被引 3

无需参考图,通过自监督学习自动评估神经渲染场景质量。

NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes without References

  • 用启发式线索和质量分数做自监督信号,不依赖人工标注。
  • 在无参考评估中平均超越17种方法超100%,媲美甚至超过有参考方法。
  • 适合需要高效、无标签质量评估的神经渲染研究者使用。

神经视图合成(NVS)如NeRF和3D高斯泼溅能从稀疏视角生成逼真场景,通常采用PSNR、SSIM、LPIPS等全参考指标评估质量。然而,这些方法受限于密集参考视图的缺乏,难以全面反映神经合成场景(NSS)的感知质量。同时,人类感知标签获取困难,导致大规模标注数据集难以构建,易引发模型过拟合与泛化能力下降。为此,我们提出NVS-SQA,一种基于自监督学习的无参考质量评估方法,无需依赖人类标签即可学习质量表示。传统自监督学习依赖“同一实例,相似表示”假设及大规模数据,但该条件在NSS质量评估中不成立。因此,我们引入启发式线索与质量分数作为学习目标,并设计专用对比对生成流程,提升学习效率与效果。实验表明,NVS-SQA在17种无参考方法中平均领先109.5%(SRCC)、98.6%(PLCC)、91.5%(KRCC),且在所有指标上均超越16种全参考方法,分别领先22.9%(SRCC)、19.1%(PLCC)、18.6%(KRCC)。

原文摘要 · Abstract (English)

Neural View Synthesis (NVS), such as NeRF and 3D Gaussian Splatting, effectively creates photorealistic scenes from sparse viewpoints, typically evaluated by quality assessment methods like PSNR, SSIM, and LPIPS. However, these full-reference methods, which compare synthesized views to reference views, may not fully capture the perceptual quality of neurally synthesized scenes (NSS), particularly due to the limited availability of dense reference views. Furthermore, the challenges in acquiring human perceptual labels hinder the creation of extensive labeled datasets, risking model overfitting and reduced generalizability. To address these issues, we propose NVS-SQA, a NSS quality assessment method to learn no-reference quality representations through self-supervision without reliance on human labels. Traditional self-supervised learning predominantly relies on the "same instance, similar representation" assumption and extensive datasets. However, given that these conditions do not apply in NSS quality assessment, we employ heuristic cues and quality scores as learning objectives, along with a specialized contrastive pair preparation process to improve the effectiveness and efficiency of learning. The results show that NVS-SQA outperforms 17 no-reference methods by a large margin (i.e., on average 109.5% in SRCC, 98.6% in PLCC, and 91.5% in KRCC over the second best) and even exceeds 16 full-reference methods across all evaluation metrics (i.e., 22.9% in SRCC, 19.1% in PLCC, and 18.6% in KRCC over the second best).

质量评估自监督神经渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。