测试视觉语言模型能否像真实施工一样一步步造房子。
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning

- 设计可迭代的建房任务,要求模型考虑结构、施工和规范约束。
- 26,000+个房屋结构经建筑标准验证,支持10项确定性结构检测。
- 适合关注生成式智能与实际工程结合的研究者。
物理世界不仅具有视觉属性,更受严格的结构与流程约束。然而当前视觉语言模型(VLM)评估仍偏重感知真实性,仅关注生成的3D布局、形状与外观是否视觉合理。现有基准极少检验模型是否理解构建物品所需的步骤与物理依赖关系,而这一能力对自动化从设计到建造的流程至关重要。为此,我们提出DreamHouse:一个针对物理生成推理的新基准,衡量模型同步满足几何、结构、可建造性与规范合规性的能力。该基准以住宅木框架建筑为场景,其工程标准完整且结果可客观验证。我们收集了超过26,000个结构,涵盖13种建筑风格,均符合施工文档标准(LOD 350),并开发了10项确定性的结构验证框架。不同于仅评估最终输出的静态基准,DreamHouse支持迭代式智能体交互:模型观察中间建造状态,生成施工动作,并接收结构化环境反馈,从而实现对规划、结构推理与自我修正能力的精细评估。对顶尖VLM的广泛实验揭示了显著的能力差距,这些差距在现有排行榜上几乎不可见。研究确立物理有效性作为与视觉真实性正交的关键评估维度,凸显物理生成推理是多模态智能中一个独特且未充分发展的前沿方向。项目地址:https://luluyuyuyang.github.io/dreamhouse
原文摘要 · Abstract (English)
The physical world is not merely visual; it is governed by rigorous structural and procedural constraints. Yet, the evaluation of vision-language models (VLMs) remains heavily skewed toward perceptual realism, prioritizing the generation of visually plausible 3D layouts, shapes, and appearances. Current benchmarks rarely test whether models grasp the step-by-step processes and physical dependencies required to actually build these artifacts, a capability essential for automating design-to-construction pipelines. To address this, we introduce DreamHouse, a novel benchmark for physical generative reasoning: the capacity to synthesize artifacts that concurrently satisfy geometric, structural, constructability, and code-compliance constraints. We ground this benchmark in residential timber-frame construction, a domain with fully codified engineering standards and objectively verifiable correctness. We curate over 26,000 structures spanning 13 architectural styles, ach verified to construction-document standards (LOD 350) and develop a deterministic 10-test structural validation framework. Unlike static benchmarks that assess only final outputs, DreamHouse supports iterative agentic interaction. Models observe intermediate build states, generate construction actions, and receive structured environmental feedback, enabling a fine-grained evaluation of planning, structural reasoning, and self-correction. Extensive experiments with state-of-the-art VLMs reveal substantial capability gaps that are largely invisible on existing leaderboards. These findings establish physical validity as a critical evaluation axis orthogonal to visual realism, highlighting physical generative reasoning as a distinct and underdeveloped frontier in multimodal intelligence. Available at https://luluyuyuyang.github.io/dreamhouse
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。