arXiv:2607.10826cs.CVcs.AI2026-07

系统评估视觉语言模型在3D缺陷检测中的表现,发现流程设计比模型本身更影响结果。

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

论文配图:3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
图 1 · 摘自论文原文
  • 构建九类细粒度缺陷的控制实验,分析渲染、提示等四个环节的影响。
  • 84种配置下超300万次评分,发现模型选择影响最大,但流程设计可改变最优方案。
  • 推荐六视角RGB为低成本高效默认方案,且需按人标注标准校准自动评判体系。

自动化评估对规模化生成式3D系统至关重要,因人工审查成本高、速度慢。然而,自动化评判的可靠性不仅取决于底层视觉语言模型(VLM),还受渲染方式、视觉输入、任务描述和参考标签构建等全流程影响。我们提出3D-DefectBench,一个用于系统分析基于VLM的3D缺陷检测流程的基准与框架。该框架补充了整体评分和成对偏好,涵盖几何、纹理和提示遵循等九类细粒度二分类缺陷,为生成器优化和评判器评估提供可操作诊断。采用平衡因子设计,我们在84种推理配置中测试了四个流程因素(VLM、相机协议、视觉输入、提示模板),完成约320万次缺陷判别,并在更广泛前沿模型上进行分阶段验证。结果显示,模型选择是与人类标签一致性的最大决定因素,但其他因素也显著影响性能,且与模型选择存在交互作用,能改变最优配置。在评估的设计空间内,紧凑的六视角RGB方案在性能上可媲美更密集的多视角设置及包含深度或法向量增强的输入,成为性价比高的默认方案。在此标准化流程下,12个VLM评判器中最优者仍落后于训练过的专业人类标注者,而当使用噪声更大的银标签替代专家共识标签时,纹理判断一致性急剧下降。这些发现表明,自动化评判应作为完整流程来评估,并在不同人类参考标准下进行校准,而非仅作为独立模型进行基准测试。我们已将标签、提示、预测和Croissant元数据发布至Hugging Face。

原文摘要 · Abstract (English)

Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the entire evaluation pipeline, not only the underlying vision-language model (VLM), but also how assets are rendered, what visual evidence is provided, how the task is specified, and how human reference labels are constructed. We introduce 3D-DefectBench, a benchmark and framework for systematic analysis of VLM-based 3D defect detection pipelines. It complements holistic ratings and pairwise preferences with nine fine-grained binary defects spanning geometry, texture, and prompt adherence, providing actionable diagnostics for generator development and judge evaluation. Using a balanced factorial design, we vary four pipeline factors, VLM, camera protocol, visual input, and prompt schema, across 84 inference designs and approximately 3.2 million scored defect decisions, followed by staged validation on a broader set of frontier models. Model choice is the largest determinant of agreement with human labels, but the remaining factors also affect performance, interact with model selection, and can change the best configuration. Within the evaluated design space, a compact six-view RGB protocol performs comparably to denser multi-view settings and inputs augmented with depth or surface normals, making it a strong cost-effective default. Under this standardized pipeline, the best of 12 VLM judges still lag behind trained human labelers, while texture agreement drops sharply when expert-consensus labels are replaced by noisier silver labels. These findings show that automated judges should be evaluated as complete pipelines and calibrated across human reference regimes, rather than benchmarked only as standalone models. We release labels, prompts, predictions, and Croissant metadata on Hugging Face.

3D生成评估基准视觉语言模型缺陷检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。