arXiv:2605.17556cs.ROcs.AI2026-05中稿 · ICRA

让机器人用视觉对齐的形态表示进行长期雕塑规划,更贴合真实艺术创作。

Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting

论文配图:Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting
图 1 · 摘自论文原文
  • 用视觉对齐的表征捕捉光照纹理,替代传统稀疏点云
  • 在三种材料上实现超过100步的长程雕塑动作,性能达顶尖水平
  • 适合需要视觉感知与长期规划的艺术类机器人任务

陶土雕塑是一项涉及灵巧操作和长期规划的精细艺术任务。我们将机器人陶土雕塑建模为形状到形状的匹配问题。以往可变形物体操作方法要么需为每个目标重新训练策略,要么依赖将状态表示为稀疏点云的动力学模型,无法有效捕捉陶土的纹理等关键特征。本文提出一种可捕获光照与纹理特征的可变形材料动力学建模方法,并支持视觉对齐的规划。在三种不同可变形材料和多种末端执行器下,我们的模型性能与当前最先进方法相当,且具备视觉规划兼容性。动作设计为单个末端执行器的参数化推压,适用于超过100步的长周期浮雕雕塑任务。最后,我们展示了视觉对齐表示在规划中的优势,也分析了其相比3D表示更具挑战性的原因。

原文摘要 · Abstract (English)

Clay sculpting is a nuanced, artistic task involving dexterous manipulation with long-horizon planning to achieve high-level goals. As a robotics problem, we formulate clay sculpting as a shape-to-shape matching challenge. Prior deformable object manipulation work either requires retraining a policy per goal or relies on dynamics models which represent state as sparse point clouds which do not capture important clay features, such as textures, well. We present a method for modeling the dynamics of deformable materials and planning for robotic sculpting in a representation that is visually-aligned, capturing lighting and texture features. With three different deformable materials and various end-effectors, we demonstrate that our dynamics model is comparable in performance to the state-of-the-art with the added benefit of being compatible with visual planning. Our actions are represented as parametrized pushes into clay with a single end-effector, which proved to be suitable for long-horizon (>100 actions) clay relief sculptures. Lastly, we show the benefits of planning in a visually-aligned representation, but also provide analysis providing evidence as to why this representation is challenging to plan in compared to 3D representations.

机器人操作视觉规划可变形物体长程控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。