用多模态模型评估3D高斯点云场景质量,兼顾视觉和结构信息。
SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

- 基于多视角图像与深度图构建结构感知的质量表征。
- 融合几何信息后,对3DGS场景的跨视角一致性评估更准确。
- 适合关注3D重建质量评估的研究者和开发者。
3D高斯点云(3DGS)已成为新视角合成与三维场景重建的有效表示方法,对可靠的质量评估需求日益增长。与传统图像质量评估(IQA)不同,3DGS场景质量不仅依赖渲染视图的感知保真度,还受空间结构和跨视角一致性等场景级因素影响。现有IQA方法受限于对2D感知线索的依赖,而通用多模态大语言模型(MLLM)未针对稳定质量回归设计,可能导致不可靠判断。为此,本文提出一种面向3DGS场景理解的多模态质量评估框架:首先,通过在VGGT编码器基础上引入专用质量头,构建3D感知的质量表征学习框架;多视角图像被编码为视图特异性特征并聚合以捕捉跨视角一致性,同时通过联合建模深度图与点云相关结构信息,融入几何线索,实现超越外观驱动特征的结构感知质量表征。其次,构建基于原始图像、深度图、点云渲染图及相机参数的接地式多模态推理机制,输入至Qwen-based MLLM进行综合判断。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。