arXiv:2501.10071cs.CVcs.MM2025-01AAAI被引 17

用语言描述质量,让点云评估更接近人眼判断。

CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment

  • 用文本描述替代分数,通过相似度匹配实现评估
  • 在多个数据集上超越现有最佳方法,尤其在主观分布预测上表现突出
  • 适合关注人眼感知与模型对齐的研究者

近年来,无参考点云质量评估(NR-PCQA)取得显著进展。然而,现有方法多采用从视觉数据到平均意见分(MOS)的直接映射,与实际主观评价机制相悖。为此,我们提出一种新型语言驱动的点云质量评估方法——CLIP-PCQA。考虑到人类更倾向于用“优秀”“差”等离散描述词表达质量而非具体分数,我们采用基于检索的映射策略,模拟主观评价过程。具体而言,借鉴CLIP思想,计算视觉特征与不同质量描述文本特征间的余弦相似度,并引入有效的对比损失和可学习提示(prompts)以增强特征提取。同时,鉴于主观实验中存在个体差异与偏差,我们将特征相似度转化为概率分布,以意见分分布(OSD)作为最终目标,而非单一MOS。实验结果表明,我们的方法在多个数据集上优于现有最先进(SOTA)方法。

原文摘要 · Abstract (English)

In recent years, No-Reference Point Cloud Quality Assessment (NR-PCQA) research has achieved significant progress. However, existing methods mostly seek a direct mapping function from visual data to the Mean Opinion Score (MOS), which is contradictory to the mechanism of practical subjective evaluation. To address this, we propose a novel language-driven PCQA method named CLIP-PCQA. Considering that human beings prefer to describe visual quality using discrete quality descriptions (e.g., "excellent" and "poor") rather than specific scores, we adopt a retrieval-based mapping strategy to simulate the process of subjective assessment. More specifically, based on the philosophy of CLIP, we calculate the cosine similarity between the visual features and multiple textual features corresponding to different quality descriptions, in which process an effective contrastive loss and learnable prompts are introduced to enhance the feature extraction. Meanwhile, given the personal limitations and bias in subjective experiments, we further covert the feature similarities into probabilities and consider the Opinion Score Distribution (OSD) rather than a single MOS as the final target. Experimental results show that our CLIP-PCQA outperforms other State-Of-The-Art (SOTA) approaches.

点云评估视觉语言模型主观对齐CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。