仅用一张照片生成可直接用于仿真的物理3D资产
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
- 基于视觉语言模型,首次实现单图生成带物理属性的3D资产
- 新3D表示减少193倍令牌数,提升生成质量且无需特殊标记
- 构建新数据集,物体类别超2倍扩展,支持机器人仿真学习
3D建模正从静态视觉表现转向可直接用于仿真与交互的物理化、可动化资产。然而,现有3D生成方法普遍忽略关键的物理与运动属性,限制了其在具身智能中的应用。为此,我们提出PhysX-Anything,首个能根据单张真实图像生成高质量、仿真就绪的3D资产的框架,包含显式几何、关节结构与物理属性。我们设计首个基于视觉语言模型的物理3D生成模型,并引入一种新型3D表示,将令牌数量减少193倍,可在标准VLM预算内实现显式几何学习,且微调时不引入特殊令牌,显著提升生成质量。此外,为克服现有物理3D数据集多样性不足的问题,我们构建了新数据集PhysX-Mobility,其物体类别比之前数据集扩展超过2倍,涵盖2000余个常见真实世界物体并附有丰富的物理标注。在PhysX-Mobility及真实图像上的大量实验表明,PhysX-Anything具备强生成性能和鲁棒泛化能力。进一步在MuJoCo风格环境中进行仿真实验,验证了所生成资产可直接用于接触丰富的机器人策略学习。我们认为PhysX-Anything可显著推动具身智能与物理仿真等下游应用的发展。
原文摘要 · Abstract (English)
3D modeling is shifting from static visual representations toward physical, articulated assets that can be directly used in simulation and interaction. However, most existing 3D generation methods overlook key physical and articulation properties, thereby limiting their utility in embodied AI. To bridge this gap, we introduce PhysX-Anything, the first simulation-ready physical 3D generative framework that, given a single in-the-wild image, produces high-quality sim-ready 3D assets with explicit geometry, articulation, and physical attributes. Specifically, we propose the first VLM-based physical 3D generative model, along with a new 3D representation that efficiently tokenizes geometry. It reduces the number of tokens by 193x, enabling explicit geometry learning within standard VLM token budgets without introducing any special tokens during fine-tuning and significantly improving generative quality. In addition, to overcome the limited diversity of existing physical 3D datasets, we construct a new dataset, PhysX-Mobility, which expands the object categories in prior physical 3D datasets by over 2x and includes more than 2K common real-world objects with rich physical annotations. Extensive experiments on PhysX-Mobility and in-the-wild images demonstrate that PhysX-Anything delivers strong generative performance and robust generalization. Furthermore, simulation-based experiments in a MuJoCo-style environment validate that our sim-ready assets can be directly used for contact-rich robotic policy learning. We believe PhysX-Anything can substantially empower a broad range of downstream applications, especially in embodied AI and physics-based simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。