arXiv:2603.23386cs.CVcs.GR2026-03International Conf…被引 2

用大模型一键拆解3D模型并生成可模拟的机械关节,提升仿真效率。

SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM

  • 通过稀疏3D向量量化自编码器,大幅减少3D token数量
  • 在PartNet-Mobility等数据集上达领先性能,支持物理仿真
  • 适合需要高质量可交互3D资产的机器人与仿真研究者

高质量可动3D资产对具身智能与物理仿真至关重要,但现有3D生成多聚焦静态网格,缺乏“可模拟”交互对象。当前方法多采用多阶段流水线,模块间误差累积严重。统一的MLLM虽可单阶段完成静态理解与可模拟资产生成,但密集体素化导致3D token序列过长、内存开销高,难以扩展至复杂可动物体。为此,我们提出SIMART,一个统一的MLLM框架,联合实现部件级分解与运动学预测。引入稀疏3D VQ-VAE,相较密集体素令牌降低70%的令牌数量,支持高保真多部件装配。SIMART在PartNet-Mobility和真实世界AIGC数据集上达到最优表现,并成功支撑物理基础的机器人仿真。

原文摘要 · Abstract (English)

High-quality articulated 3D assets are indispensable for embodied AI and physical simulation, yet 3D generation still focuses on static meshes, leaving a gap in "sim-ready" interactive objects. Most recent articulated object creation methods rely on multi-stage pipelines that accumulate errors across decoupled modules. Alternatively, unified MLLMs offer a single-stage path to joint static asset understanding and sim-ready asset generation. However dense voxel-based 3D tokenization yields long 3D token sequences and high memory overhead, limiting scalability to complex articulated objects. To address this, we propose SIMART, a unified MLLM framework that jointly performs part-level decomposition and kinematic prediction. By introducing a Sparse 3D VQ-VAE, SIMART reduces token counts by 70% vs. dense voxel tokens, enabling high-fidelity multi-part assemblies. SIMART achieves state-of-the-art performance on PartNet-Mobility and in-the-wild AIGC datasets, and enables physics-based robotic simulation.

3D生成可动建模物理仿真大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。