用物理逻辑生成可交互3D资产,让虚拟世界更真实
PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

- 先规划物理蓝图,再用扩散模型生成带精确运动参数的3D资产
- 基于15万条带四层物理标注的数据集,生成结果可直接用于仿真
- 适合虚拟现实、智能体训练等需要真实物理交互的场景
生成具备物理特性的3D资产是互动虚拟世界和具身AI的关键瓶颈。现有方法多关注静态几何,忽视交互所需的功能属性。我们提出,交互资产生成必须基于功能逻辑与分层物理。为此,我们推出PhysForge,一个由PhysDB(包含15万条资产、四层物理标注)支持的解耦两阶段框架。首先,视觉语言模型作为“物理建筑师”,规划出定义材料、功能与运动约束的“分层物理蓝图”。其次,基于物理的扩散模型通过新颖的KineVoxel注入(KVI)机制,合成高保真几何与精确运动参数。实验表明,PhysForge生成的资产功能合理、可直接用于仿真,为交互式3D内容与具身智能体提供可靠数据引擎。
原文摘要 · Abstract (English)
Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We propose that interactive asset generation must be rooted in functional logic and hierarchical physics. To bridge this gap, we introduce PhysForge, a decoupled two-stage framework supported by PhysDB, a large-scale dataset of 150,000 assets with four-tier physical annotations. First, a VLM acts as a "physical architect" to plan a "Hierarchical Physical Blueprint" defining material, functional, and kinematic constraints. Second, a physics-grounded diffusion model realizes this blueprint by synthesizing high-fidelity geometry alongside precise kinematic parameters via a novel KineVoxel Injection (KVI) mechanism. Experiments demonstrate that PhysForge produces functionally plausible, simulation-ready assets, providing a robust data engine for interactive 3D content and embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。