arXiv:2504.08125cs.CV2025-04CVPR被引 10

用视觉大模型自动评估3D生成质量,更贴近真人判断。

Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects

  • 用微调的视觉大模型分析3D表面法向,评估外观与质量。
  • 在用户偏好测试中优于现有无参考评估方法。
  • 无需真实对比数据,适合研究文本生成3D的新方向。

文本到3D生成技术的快速进展亟需高效且可靠的评估指标,以贴近人类判断。现有指标如PSNR和CLIP依赖真实数据或仅关注提示匹配度,难以满足需求。为此,我们提出Gen3DEval,一种基于专门微调的视觉大语言模型(vLLM)的新型评估框架。该框架通过分析3D表面法向,无需真实参照即可评估文本一致性、外观表现和表面质量,有效弥合自动化评估与用户偏好之间的差距。在用户对齐评估中,Gen3DEval相较于最先进的无任务特异性模型表现更优,展现出全面且可访问的潜力,有望成为未来文本到3D生成研究的基准工具。

原文摘要 · Abstract (English)

Rapid advancements in text-to-3D generation require robust and scalable evaluation metrics that align closely with human judgment, a need unmet by current metrics such as PSNR and CLIP, which require ground-truth data or focus only on prompt fidelity. To address this, we introduce Gen3DEval, a novel evaluation framework that leverages vision large language models (vLLMs) specifically fine-tuned for 3D object quality assessment. Gen3DEval evaluates text fidelity, appearance, and surface quality by analyzing 3D surface normals, without requiring ground-truth comparisons, bridging the gap between automated metrics and user preferences. Compared to state-of-the-art task-agnostic models, Gen3DEval demonstrates superior performance in user-aligned evaluations, placing it as a comprehensive and accessible benchmark for future research on text-to-3D generation. The project page can be found here: \href{https://shalini-maiti.github.io/gen3deval.github.io/}{https://shalini-maiti.github.io/gen3deval.github.io/}.

3D生成评估方法视觉大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。