arXiv:2606.18451cs.LG2026-06被引 1

提出可复现的视觉语言模型评分协议,验证现有方法无法准确评估单图生成3D网格质量。

A Cross-Model VLM-Judge Protocol for Single-Image 3D Mesh Quality (and Why Cheap Proxies Fall Short)

论文配图:A Cross-Model VLM-Judge Protocol for Single-Image 3D Mesh Quality (and Why Cheap Proxies Fall Short)
图 1 · 摘自论文原文
  • 设计24视角无头渲染+双视觉语言模型评判,通过一致性校验确保结果可靠
  • 几何有效性仅在缺陷明显时有效,整体相关性弱;渲染相似度基本随机
  • 揭示廉价代理的局限性,推荐使用新协议作为标准评估工具

单图生成3D网格技术快速进步,但缺乏无需人工参与的质量评估标准。研究者常依赖低成本自动代理(如渲染空间CLIP相似度和网格几何有效性统计),但其与真实感知质量的相关性未被验证。本文提出并验证了一种可复现的VLM-Judge评估协议:固定24视角无头渲染装置、两个独立的视觉语言模型评判家族,以及强制位置偏差校正机制(同时查询两种展示顺序并仅保留一致结论)。两个评判家族间一致性较高(Cohen's kappa = 0.66),远超偶然水平。以该协议为基准,发现几何有效性平均相关性弱(因呈双峰分布),且低于预注册目标;渲染CLIP则处于随机水平。学习得到的Bradley-Terry头最终坍缩至单一manifoldness统计量,使渲染CLIP获得负权重,说明学习特征权重无效。代理本身也呈双峰分布:在可见几何缺陷对比中显著优于随机,在模糊对比中仍为随机,表明其仅在缺陷显著时追踪评判结果。因此,建议在测试条件下(两个前馈生成器、Google Scanned Objects数据集、面部缺失退化方案)采用该VLM-judge协议作为可靠评估方式,并反对将几何或CLIP代理用作优化目标。

原文摘要 · Abstract (English)

Single-image-to-3D generators are improving quickly, but there is no agreed, human-free way to tell whether one generated mesh is better than another. Practitioners commonly rely on cheap automatic proxies (render-space CLIP similarity and mesh geometry-validity statistics), yet how well these track perceived quality is unestablished. We make two contributions. First, we propose and validate a reproducible VLM-judge evaluation protocol: a fixed 24-view headless render rig, two independent vision-language judge families, and a mandatory position-bias correction that queries both presentation orders and keeps only order-consistent verdicts. The two judge families agree substantially with each other (Cohen's kappa = 0.66), well above the chance-agreement floor. Second, using this protocol as the reference, we show the cheap proxies do not substitute for it. Geometry validity is only a weak signal on average (because, as we show, it is bimodal) and stays below our pre-registered target, while render-CLIP is at chance. A learned Bradley-Terry head collapses onto a single manifoldness statistic (giving render-CLIP a negative weight) and matches geometry-only exactly, so learning the feature weights buys nothing. The proxy is also bimodal: it is significantly above chance on contrasts with visible geometric defects but at chance on ambiguous contrasts, consistent with geometry validity tracking the judge only when the defect is visually salient. We therefore recommend the VLM-judge protocol as a reliable, reproducible evaluator under the conditions tested (two feed-forward generators on Google Scanned Objects, with a face-drop degradation regime) and advise against geometry/CLIP proxies as optimization targets.

3D生成质量评估视觉语言模型网格评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。