arXiv:2512.24428cs.RO2025-12被引 1

一秒钟生成机器人可用的3D模型,支持实时抓取与避障。

Subsecond 3D Mesh Generation for Robot Manipulation

  • 用RGB-D图像端到端生成高精度3D网格,速度<1秒/对象。
  • 结合语义分割、扩散生成与点云配准,实现上下文精准建模。
  • 适合需要快速感知与规划的机器人操作场景。

3D网格是计算机科学与工程中的基础表示形式,在机器人领域尤为重要,因其能直接反映物体与物理世界的交互特性,支撑稳定抓取预测、碰撞检测与动力学仿真等核心能力。尽管近年自动3D网格生成方法取得进展,有望实现机器人实时感知,但仍面临两大挑战:一是高质量网格生成速度过慢,通常需数十秒;二是生成本身不足,网格必须在场景中正确分割,并准确对齐尺度与位姿,否则将引入新瓶颈。本文提出一个端到端系统,仅需单张RGB-D图像即可在1秒内生成高质量、上下文精确的3D网格。该流程整合开放词汇语义分割、加速的扩散式网格生成和鲁棒点云配准,各模块均优化速度与精度。我们在真实抓取任务中验证其有效性,证明该方法可作为机器人感知与规划的实用、按需3D表示。

原文摘要 · Abstract (English)

3D meshes are a fundamental representation widely used in computer science and engineering. In robotics, they are particularly valuable because they capture objects in a form that aligns directly with how robots interact with the physical world, enabling core capabilities such as predicting stable grasps, detecting collisions, and simulating dynamics. Although automatic 3D mesh generation methods have shown promising progress in recent years, potentially offering a path toward real-time robot perception, two critical challenges remain. First, generating high-fidelity meshes is prohibitively slow for real-time use, often requiring tens of seconds per object. Second, mesh generation by itself is insufficient. In robotics, a mesh must be contextually grounded, i.e., correctly segmented from the scene and registered with the proper scale and pose. Additionally, unless these contextual grounding steps remain efficient, they simply introduce new bottlenecks. In this work, we introduce an end-to-end system that addresses these challenges, producing a high-quality, contextually grounded 3D mesh from a single RGB-D image in under one second. Our pipeline integrates open-vocabulary object segmentation, accelerated diffusion-based mesh generation, and robust point cloud registration, each optimized for both speed and accuracy. We demonstrate its effectiveness in a real-world manipulation task, showing that it enables meshes to be used as a practical, on-demand representation for robotics perception and planning.

3D生成机器人感知实时建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。