用多模态模型自动评估点云质量,无需参考图。
PIT-QMM: A Large Multimodal Model For No-Reference Point Cloud Quality Assessment
- 融合文本、图像和点云数据端到端预测质量得分
- 在主流基准上超越现有方法,训练迭代更少
- 可定位并识别失真区域,提升模型可解释性
大型多模态模型(LMMs)在图像和视频质量评估领域取得显著进展,但在3D资产领域尚未充分探索。本文关注无参考点云质量评估(NR-PCQA),即在无参考的情况下自动评估点云的感知质量。我们观察到,文本描述、2D投影和3D点云视图等不同模态的数据能提供互补的质量信息。基于此,我们构建了PIT-QMM,一种新型的用于NR-PCQA的大型多模态模型,可端到端处理文本、图像和点云数据以预测质量分数。大量实验表明,该方法在主流基准上显著优于现有最先进方法,且所需训练迭代次数更少。此外,我们还展示了该框架可实现失真定位与识别,为模型可解释性和交互性开辟新路径。代码与数据集见 https://www.github.com/shngt/pit-qmm。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have recently enabled considerable advances in the realm of image and video quality assessment, but this progress has yet to be fully explored in the domain of 3D assets. We are interested in using these models to conduct No-Reference Point Cloud Quality Assessment (NR-PCQA), where the aim is to automatically evaluate the perceptual quality of a point cloud in absence of a reference. We begin with the observation that different modalities of data - text descriptions, 2D projections, and 3D point cloud views - provide complementary information about point cloud quality. We then construct PIT-QMM, a novel LMM for NR-PCQA that is capable of consuming text, images and point clouds end-to-end to predict quality scores. Extensive experimentation shows that our proposed method outperforms the state-of-the-art by significant margins on popular benchmarks with fewer training iterations. We also demonstrate that our framework enables distortion localization and identification, which paves a new way forward for model explainability and interactivity. Code and datasets are available at https://www.github.com/shngt/pit-qmm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。