揭示视觉语言动作模型在接触密集任务中失败原因并提出有效解决方案
Demystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them

- 发现两类失败模式:精度错配与力信号结构特殊性
- 新方法FACT在5个任务上平均成功率提升至66%(前最优41%)
- 适合研究机器人操控、具身智能的学者和工程师参考
我们研究视觉-语言-动作模型在需要精确物理交互的接触密集型操作任务中表现不佳的原因。以往工作主要通过力增强架构和训练正则化来缓解接触失败,但根本原因仍不明确。我们识别出两种不同的失败模式:精度失败源于流匹配策略训练不一致,力失败则由力信号的独特结构引起。针对每种模式设计针对性修复机制,并集成到FACT系统中。在近2500次真实世界轨迹评估中,FACT在五个接触密集任务上的平均成功率达到66%,显著优于最佳基线的41%。
原文摘要 · Abstract (English)
We address the problem of understanding when and why Vision-Language-Action models struggle with contact-rich manipulation tasks that require precise physical interaction. Prior work has primarily focused on addressing contact failures through force-augmented architectures and training-time regularizers, yet the root causes of these failures remain underexplored. We identify two distinct failure modes underlying this gap. Precision failures are rooted in a flow-matching policy training mismatch, and force failures arise from the distinctive structure of force signals. We address each failure mode with a targeted mechanism and combine them into FACT, which achieves 66% average success rate across five contact-rich tasks against 41% for the best prior baseline, in an evaluation spanning almost 2,500 real-world rollouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。