arXiv:2608.18627cs.CV2026-08

用强化学习提升点云质量评估的跨数据集泛化能力

PCQA-R1: Advancing Generalized 3D Point Cloud Quality Assessment with Reinforcement Learning

论文配图:PCQA-R1: Advancing Generalized 3D Point Cloud Quality Assessment with Reinforcement Learning
图 1 · 摘自论文原文
  • 基于强化学习构建推理链数据集,实现冷启动训练
  • 在5个基准上达成顶尖跨域泛化性能,保持域内精度
  • 通过高斯近邻奖励避免评分漂移,适合标注稀缺场景

无参考点云质量评估(PCQA)近年受到关注,用于衡量和优化点云视觉体验。然而,大型多模态模型(LMM)在此领域应用较少。以往基于LMM的方法主要依赖监督微调直接预测数值质量得分,难以在不同主观评分尺度和标注有限的数据集间泛化。关键难点在于绝对主观评分回归在不同数据集间易失效,而相对质量排序更具稳定性。本文提出PCQA-R1,首个用于3D点云质量评估的强化学习型LMM,同时建模质量理解与评分。基于组相对策略优化(GRPO)策略,我们构建了推理链数据集PCQA-CoT,通过逆向推理策略训练LMM生成推理过程作为冷启动数据。进一步引入高斯邻近奖励,通过锚定源数据主观评分范围防止校准漂移。实验表明,PCQA-R1在五个基准上实现顶尖跨数据集泛化性能,且域内精度具竞争力。消融实验证明排序机制、高斯奖励及冷启动轨迹的关键作用。

原文摘要 · Abstract (English)

No-reference point cloud quality assessment (PCQA) has been an active topic in recent years and is used to measure and optimize the visual experience of point clouds. However, large multimodal models (LMMs) have rarely been explored in this area. Previous LMM-based methods mainly rely on supervised fine-tuning to directly predict numerical quality scores, lacking the ability to generalize across datasets with heterogeneous MOS scales and limited annotations. A key difficulty is that absolute MOS regression can be brittle across datasets with different score scales and distortion distributions, whereas relative quality ranking is more stable under such shifts. In this paper, we present PCQA-R1, the first reinforcement learning LMM for 3D point cloud quality assessment to simultaneously model quality understanding and scoring. Built upon the group relative policy optimization (GRPO) strategy, PCQA-R1 first constructs a chain-of-thought dataset, PCQA-CoT, which serves as cold-start training data through a reverse reasoning strategy that teaches the LMM to generate its reasoning process. We further introduce a Gaussian proximity reward that prevents calibration drift by anchoring score predictions to the source MOS range. Experimental results demonstrate that PCQA-R1 achieves state-of-the-art cross-dataset generalization across five benchmarks and competitive in-domain accuracy. Ablation studies support the role of ranking, Gaussian reward, and cold-start traces.

点云评估强化学习多模态模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。