首个支持缺陷诊断与可解释评估的点云质量评测基准。
PointQ-Bench: Benchmarking Diagnostic and Interpretable Point Cloud Quality Assessment

- 构建包含3083个点云的多类型缺陷数据集,支持缺陷定位与描述。
- 发现现有模型在细粒度诊断与质量校准上表现不佳,存在感知-诊断差距。
- 适用于需要可解释性输出的工业质检、3D重建与生成内容评估场景。
点云质量在3D采集、重建、渲染与感知中至关重要,但现有点云质量评估(PCQA)研究仍以标量分数预测为主。实际检测中,质量评估需识别缺陷、分类问题类型、评估下游可用性并提供证据支撑的描述,而当前基准未覆盖这些需求。我们提出PointQ-Bench,一个从标量评分扩展至全面质量理解的基准。该数据集包含3,083个点云,涵盖真实扫描、模拟失真和AI生成内容,覆盖八类主要问题。每个样本标注了平均意见得分(MOS)、质量等级、问题标签、专家级描述及12,332个问答对。支持异常感知、缺陷诊断、可用性评分等感知任务,以及开放式的质量报告认知任务。为评估自由文本描述,我们提出SSFRQ-5D五维评估协议,并通过人机一致性分析验证。14个视觉语言模型与传统基线的实验表明:当前模型虽具粗粒度缺陷感知能力,但在基于证据的诊断与质量校准上表现薄弱;强2D多模态大模型普遍优于现有3D视觉语言模型,额外视图或点级输入的增益不一致,尤其在边界模糊条件下差异显著。总体而言,PointQ-Bench为提升可靠且可解释的点云质量理解提供了诊断测试平台。
原文摘要 · Abstract (English)
Point cloud quality plays a critical role in 3D acquisition, reconstruction, rendering, and perception, yet existing point cloud quality assessment (PCQA) research remains largely centered on scalar score prediction. In practical inspection scenarios, quality assessment often involves identifying defects, characterizing dominant issue types, assessing downstream usability, and providing evidence-supported descriptions, which are not explicitly evaluated by current benchmarks. We introduce PointQ-Bench, a benchmark designed to extend PCQA from scalar scoring toward comprehensive quality understanding. PointQ-Bench consists of 3,083 point clouds spanning authentic scans, simulated distortions, and AI-generated content, covering eight major issue types. Each sample is annotated with mean opinion scores (MOS), quality levels, issue tags, expert-grounded descriptions, and 12,332 question-answer pairs. The benchmark supports three perception-oriented tasks: anomaly sensing, defect diagnosis, and usability grading, as well as a cognition-oriented task of open-ended quality reporting. To evaluate free-form quality descriptions, we further propose SSFRQ-5D, a five-dimensional evaluation protocol validated through human-AI agreement analysis. Extensive experiments on 14 vision-language models and traditional PCQA baselines reveal a consistent perception-diagnosis gap: while current models exhibit emerging abilities in coarse defect perception, they struggle with grounded diagnosis and quality calibration. Strong 2D MLLMs generally outperform existing 3D VLMs, and the benefit of additional views or point-level inputs is non-uniform, varying across tasks, data sources, and models, particularly under boundary-ambiguous conditions. Overall, PointQ-Bench provides a diagnostic testbed for advancing reliable and interpretable point cloud quality understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。