提出分频控制的力感知机器人操作模型,提升接触任务反应速度与成功率。
FAVLA: A Force-Adaptive Fast-Slow VLA model for Contact-Rich Robotic Manipulation
- 慢速视觉语言模型与快速动作执行器解耦,分别处理感知与控制
- 在接触力较小时成功率提升32%,响应延迟降低40%
- 根据预测力变化动态调节执行频率,适合高精度接触操作场景
力/力矩反馈能显著提升视觉-语言-动作(VLA)模型在接触密集型操作任务中的表现,但现有方法常在单一频率下融合多模态信息,忽略真实机器人传感器采样率差异,导致高频接触信号被迫降采样,影响实时修正。结合常见的视觉语言模型-动作专家(VLM-AE)流水线中动作块在昂贵的VLM更新之间大量开环执行,统一频率融合常造成对冲击、粘滑和力突变的响应延迟。本文提出FAVLA,一种力自适应的快-慢VLA模型,将慢速感知规划与快速接触感知控制分离。慢速VLM以固定低频运行,编码多模态信息并预测近未来力变化;快速动作专家(AE)以可变高频执行,基于最新力序列数据生成响应式动作。我们进一步引入力适配器,将高频力特征注入到多个AE层,并根据VLM预测的力变化自适应调度AE执行频率。大量接触密集型任务实验表明,FAVLA显著优于基线,在较小接触力下仍保持更高反应性与成功率。
原文摘要 · Abstract (English)
Force/torque feedback can substantially improve Vision-Language-Action (VLA) models on contact-rich manipulation, but most existing approaches fuse all modalities at a single operating frequency. This design ignores the mismatched sampling rates of real robot sensors, forcing downsampling of the high-frequency contact cues needed for reactive correction. Combined with common VLM-action-expert (AE) pipelines that execute action chunks largely open loop between expensive VLM updates, unified-frequency fusion often yields delayed responses to impacts, stick-slip, and force spikes. We propose FAVLA, a force-adaptive fast-slow VLA that decouples slow perception planning from fast contact-aware control. FAVLA runs a slow VLM at a fixed low frequency to encode modalities to produce latent representations and to predict near-future force variation. A fast AE then executes at a variable high frequency, conditioning on the latest force sequence data to generate reactive actions. We further introduce a force adapter that injects high-frequency force features into multiple AE layers, and adaptively schedules the AE's execution frequency based on the VLM's predicted force variation. Extensive experiments on contact-rich tasks demonstrate that FAVLA significantly outperforms baselines, achieving superior reactivity and success rates, especially with a smaller contact force during manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。