评测大模型生成可执行参数化3D程序的能力,揭示其结构理解短板。
P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

- 构建统一协议的基准,评估文本/图像到3D及装配体生成能力。
- 模型能保整体形状但难还原精确参数几何与部件数量。
- 适合关注3D生成结构一致性与参数精度的研究者使用。
多模态大语言模型可编写代码生成复杂程序,并用程序进行3D建模,为基于其先验知识与推理能力的3D生成开辟新路径。然而现有基准很少通过代码评估3D建模。此类建模不仅需可运行代码,更要求从文本或视觉输入生成几何精确、语义对齐且装配一致的参数化3D程序。我们提出P3D-Bench,一个面向参数化3D生成的基准。不同于3D网格,参数化3D程序显式暴露尺寸、构造操作与部件关系,可揭示模型是否恢复设计结构而非仅外观。在统一协议下,该基准涵盖三类任务(文本到3D、图像到3D、装配体3D),并评分输出的可执行性、几何保真度、拓扑正确性、文本引导约束、多视角语义对齐及部件级结构。我们在400个文本案例、400个图像案例和203个标注装配体上评估前沿多模态与纯文本大模型,以领域专用模型为参照。结果发现:装配体任务最困难,模型仍难以将多个部件组合成连贯结构;模型常能恢复目标物体的整体形状与语义身份,却无法复现输入指定的精确参数几何;部件级建模在装配体中表现薄弱,既无法还原各部件几何,也难以确定正确部件数量。这些结果确立了P3D-Bench作为评估参数化3D生成中精确几何与部件级结构的基准地位。
原文摘要 · Abstract (English)
Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling, which opens up a new avenue for 3D generation powered by their priors, world knowledge and reasoning. Yet existing benchmarks rarely evaluate 3D modeling through code. Such modeling demands more than runnable code: from a text or visual specification, a model must generate a parametric 3D program that is geometrically precise, semantically aligned and assembly-consistent. We introduce P3D-Bench, a benchmark for parametric 3D generation. Unlike a 3D mesh, a parametric 3D program exposes explicit dimensions, construction operations and part relations, revealing whether a model recovers a design's structure, not just its appearance. Under a unified protocol, P3D-Bench covers three task families (Text-to-3D, Image-to-3D and Assembly-3D) and scores each output for executability, geometric fidelity, topology, text-grounded constraints, multiview semantic alignment and part-level structure. We evaluate frontier MLLMs and text-only LLMs on 400 text cases, 400 image cases and 203 annotated assemblies, with domain-specific models as reference points. Our extensive evaluation yields three findings. First, assemblies are the hardest setting, where models still fail to compose multiple parts into a coherent structure. Second, models can often recover the global shape and semantic identity of the target object, yet fail to reproduce the precise parametric geometry specified by the input. Third, part-level modeling remains weak on assemblies, where models recover neither the geometry of each part nor the right number of parts. These results position P3D-Bench as a benchmark for evaluating precise parametric geometry and part-level structure in parametric 3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。