arXiv:2512.23864cs.ROcs.CV2025-12被引 16

让机器人通过预测触觉未来,学会感知物理接触。

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

  • 用高分辨率触觉图像+多视角视觉构建层级感知系统
  • 通过预测未来触觉信号,实现95%的接触任务成功率
  • 适合需要精细触觉交互的机器人控制研究者

视觉-语言-动作(VLA)模型虽能泛化至机器人控制,却缺乏对物理接触的感知。为此,我们提出DreamTacVLA,通过学习“预知未来触觉”来赋予模型接触物理理解能力。该模型采用分层感知架构:高分辨率触觉图像作为微观视觉输入,结合腕部摄像头局部视觉与第三人称宏观视觉。为融合多尺度感知,先用分层空间对齐(HSA)损失训练统一策略,使触觉特征与腕部及第三人称视图的空间对应;再通过触觉世界模型微调,预测未来触觉信号以深化对细粒度接触动态的理解。为缓解触觉数据稀缺与传感器易损问题,构建了由高保真数字孪生和真实实验组成的混合大规模数据集。通过预判未来触觉状态,DreamTacVLA建立丰富的接触物理模型,并基于实时观测与想象后果决策。在多种接触密集型操作任务中,其表现超越现有VLA基线,最高达95%成功率达,凸显理解物理接触对构建鲁棒触觉智能体的重要性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain blind to physical contact. Consequently, they struggle with contact-rich manipulation tasks that require reasoning about force, texture, and slip. While some approaches incorporate low-dimensional tactile signals, they fail to capture the high-resolution dynamics essential for such interactions. To address this limitation, we introduce DreamTacVLA, a framework that grounds VLA models in contact physics by learning to feel the future. Our model adopts a hierarchical perception scheme in which high-resolution tactile images serve as micro-vision inputs coupled with wrist-camera local vision and third-person macro vision. To reconcile these multi-scale sensory streams, we first train a unified policy with a Hierarchical Spatial Alignment (HSA) loss that aligns tactile tokens with their spatial counterparts in the wrist and third-person views. To further deepen the model's understanding of fine-grained contact dynamics, we finetune the system with a tactile world model that predicts future tactile signals. To mitigate tactile data scarcity and the wear-prone nature of tactile sensors, we construct a hybrid large-scale dataset sourced from both high-fidelity digital twin and real-world experiments. By anticipating upcoming tactile states, DreamTacVLA acquires a rich model of contact physics and conditions its actions on both real observations and imagined consequences. Across contact-rich manipulation tasks, it outperforms state-of-the-art VLA baselines, achieving up to 95% success, highlighting the importance of understanding physical contact for robust, touch-aware robotic agents.

机器人控制触觉感知多模态学习世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。