arXiv:2501.13507cs.RO2025-01被引 1

用视觉语言模型和傅里叶形状表示,实现多颗粒物的自主搬运与塑形。

Iterative Shaping of Multi-Particle Aggregates based on Action Trees and VLM

  • 结合视觉语言模型规划抓取与推移动作,实现高层任务决策。
  • 通过傅里叶级数参数化颗粒轮廓,动态生成轨迹点保持群体凝聚力。
  • 适用于需要精准操控散颗粒系统的场景,如材料合成与柔性制造。

本文研究双臂机器人对多颗粒物聚集体的操控问题。所提方法通过一系列塑形与推移动作,实现分散颗粒的自主运输。该任务面临两大挑战:高层任务规划与轨迹执行。在任务规划方面,利用视觉语言模型(VLM)支持工具抓取、非握持式颗粒推移等基础动作;在轨迹执行方面,采用截断傅里叶级数表示颗粒聚集体的轮廓,实现闭合形状的高效参数化。通过自适应计算轨迹航点,综合考虑群体凝聚力与聚集体几何质心,以应对空间分布与整体运动变化。真实世界实验表明,该方法在主动塑造与操控多颗粒聚集体的同时,能有效维持系统高凝聚力。

原文摘要 · Abstract (English)

In this paper, we address the problem of manipulating multi-particle aggregates using a bimanual robotic system. Our approach enables the autonomous transport of dispersed particles through a series of shaping and pushing actions using robotically-controlled tools. Achieving this advanced manipulation capability presents two key challenges: high-level task planning and trajectory execution. For task planning, we leverage Vision Language Models (VLMs) to enable primitive actions such as tool affordance grasping and non-prehensile particle pushing. For trajectory execution, we represent the evolving particle aggregate's contour using truncated Fourier series, providing efficient parametrization of its closed shape. We adaptively compute trajectory waypoints based on group cohesion and the geometric centroid of the aggregate, accounting for its spatial distribution and collective motion. Through real-world experiments, we demonstrate the effectiveness of our methodology in actively shaping and manipulating multi-particle aggregates while maintaining high system cohesion.

机器人操控视觉语言模型颗粒系统轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。