用真实照片生成10万张可模拟的桌面场景,让机器人训练更贴近现实。
TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

- 从真实网络图片自动重建高精度桌面布局,不靠虚构
- 构建包含10万张物理一致场景与无碰撞操作轨迹的数据集
- 适合做泛化机器人抓取与复杂环境交互的研究者
通用机器人操作策略的发展受限于大规模、高保真场景数据的缺乏。现有自动化合成方法依赖文本生成布局或简化程序化构造,常导致物理不合理且无法捕捉真实人类环境中的密集杂乱。本文提出TableVerse,一种完全自动化的Real2Sim流程,将布局生成范式从想象转向基于无结构野外图像的确定性重建。该框架可无缝处理非脚本化的互联网媒体,生成带精确尺度、真实拓扑和机械稳定性验证的仿真桌面环境。同时集成任务条件轨迹生成框架,自动生成高质量、无碰撞的拾取-放置示范。基于此完整流程,我们构建了TableVerse-100K数据集,包含10万个唯一、物理一致的环境及其交互操作轨迹。通过覆盖多样的物品组合、真实的空间分布与高质量演示,TableVerse-100K为未来泛化机器人操作研究提供了高度可扩展且高保真的数据基础。
原文摘要 · Abstract (English)
The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shifts the paradigm from imaginative layout generation to deterministic reconstruction from unstructured, in-the-wild image data. Our framework seamlessly processes unscripted internet media into high-fidelity, simulation-ready tabletop environments with accurate metric scales, authentic topologies, and verified mechanical stability. Furthermore, an automated task-conditioned trajectory generation framework is integrated to synthesize high-quality, collision-free pick-and-place demonstrations. Leveraging this complete pipeline, we construct the TableVerse-100K Dataset, a large-scale corpus comprising 100,000 unique, physically consistent environments paired with interactive manipulation trajectories. By capturing diverse asset compositions, realistic spatial distributions, and high-quality demonstrations, TableVerse-100K establishes a highly scalable and high-fidelity data foundation, providing significant value to facilitate future research in generalizable robotic manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。