arXiv:2505.07600cs.ROcs.CV2025-05中稿 · ICRA

让视觉语言模型学会看衣服折叠的动态过程,提升操作准确性

Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding

  • 用时序上下文隐式建模衣物状态,避免手动定义复杂形态
  • 在褶皱或操作失败场景下,动作成功率显著提升
  • 适合做具身智能中复杂柔性物体操作的研究者参考

由于衣物具有复杂的动力学特性、高度可变形性以及频繁的自遮挡,对其进行操作极具挑战。衣物几乎存在无限多的构型,难以显式定义状态表示。本文分析了 BiFold 模型,该模型从视觉观测中预测语言引导的抓取与放置动作,同时通过端到端学习隐式编码衣物状态。为应对褶皱衣物或操作失败后的恢复场景,BiFold 利用时序上下文改进状态估计。我们分析了模型内部表征,发现其微调和时序上下文能有效实现文本与图像区域的对齐,以及时间上的一致性。

原文摘要 · Abstract (English)

Manipulating clothes is challenging due to their complex dynamics, high deformability, and frequent self-occlusions. Garments exhibit a nearly infinite number of configurations, making explicit state representations difficult to define. In this paper, we analyze BiFold, a model that predicts language-conditioned pick-and-place actions from visual observations, while implicitly encoding garment state through end-to-end learning. To address scenarios such as crumpled garments or recovery from failed manipulations, BiFold leverages temporal context to improve state estimation. We examine the internal representations of the model and present evidence that its fine-tuning and temporal context enable effective alignment between text and image regions, as well as temporal consistency.

视觉语言模型柔性物体操作时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。