arXiv:2607.18231cs.RO2026-07

用力觉历史增强视觉语言模型,让机器人更懂重复接触动作

FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

论文配图:FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
图 1 · 摘自论文原文
  • 用变分自编码器将力觉数据压缩成记忆令牌
  • 在三个重复操作任务中达80%以上成功率
  • 适合需要连续触觉交互的机械臂控制场景

视觉-语言-动作(VLA)模型在机器人操作中表现出色,近期的带记忆VLA通过依赖过去图像或语言摘要放松了马尔可夫假设。基于视觉的记忆方法依赖采样的历史图像帧,但计算开销大,且在视觉变化微小的场景下表现受限,如多次按压按钮。本文提出FM-VLA,一种基于力觉记忆的VLA模型,支持非马尔可夫、高接触密度的操作任务中的时序推理。通过预训练变分自编码器(VAE)重构力信号时间序列,将力觉历史编码为紧凑的力觉记忆令牌,并将其与短期状态历史一起作为额外条件输入动作专家模块,使模型能利用累积的接触事件历史指导操作。在三个依赖记忆的任务上评估:寻找隐藏方块、按按钮、擦拭盘子指定次数。轻量级力觉记忆在几乎无额外推理开销下达到超过80%的成功率,显著优于基线方法。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/

机器人操作力觉感知记忆机制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。