用单图自动生成10万套可模拟的3D物体资产
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
- 单图转3D:自动构建仿真可用且带语义标注的数字孪生体
- 建成10万件高质量资产,含物理属性与功能标注
- 适合机器人抓取、场景合成与视觉问答数据生成
模拟学习为提升机器人操作能力提供了有效基础,但常受限于缺乏足够规模和多样性的数据生成可用数字资产。本文提出ManiTwin,一个自动化高效的数据生成就绪数字物体生成管道。该管道将单张图像转化为可直接用于仿真的语义标注3D资产,支持大规模机器人操作数据生成。基于此,我们构建了包含10万件高质量标注3D资产的ManiTwin-100K数据集。每个资产均配备物理属性、语言描述、功能标注及经验证的操作建议。实验表明,ManiTwin实现了高效的资产合成与标注流程,所建数据集具备高质量与多样性,适用于操作数据生成、随机场景合成及视觉问答数据构建,为可扩展的仿真数据合成与策略学习奠定坚实基础。
原文摘要 · Abstract (English)
Learning in simulation provides a useful foundation for scaling robotic manipulation capabilities. However, this paradigm often suffers from a lack of data-generation-ready digital assets, in both scale and diversity. In this work, we present ManiTwin, an automated and efficient pipeline for generating data-generation-ready digital object twins. Our pipeline transforms a single image into simulation-ready and semantically annotated 3D asset, enabling large-scale robotic manipulation data generation. Using this pipeline, we construct ManiTwin-100K, a dataset containing 100K high-quality annotated 3D assets. Each asset is equipped with physical properties, language descriptions, functional annotations, and verified manipulation proposals. Experiments demonstrate that ManiTwin provides an efficient asset synthesis and annotation workflow, and that ManiTwin-100K offers high-quality and diverse assets for manipulation data generation, random scene synthesis, and VQA data generation, establishing a strong foundation for scalable simulation data synthesis and policy learning. Our webpage is available at https://manitwin.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。