arXiv:2607.23782cs.RO2026-07被引 3

首个大规模触觉感知的多模态机器人模型,实现精细操作与离线优化。

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

论文配图:$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
图 1 · 摘自论文原文
  • 融合视觉、触觉与语言,通过分阶段训练集成触觉路径。
  • 在真实机器人任务中成功率超基线,模拟任务平均达63.8%。
  • 适合需要精细触觉反馈的机器人操控场景,如柔性物体操作。

我们提出 $N_0$-VTLA,一种视觉-触觉-语言-动作(VTLA)基础模型,具备(1)基于触觉感知与反馈控制的精细接触式操作能力,(2)从存储部署数据中进行离线策略优化。基于现有视觉骨干网络,我们设计了包含视觉-触觉预训练、分阶段触觉路径整合及优势条件化离线策略优化的训练方法。预训练阶段,策略从 NeoData——我们构建的大规模视觉-触觉机器人数据集——中学习广泛的接触先验;据我们所知,$N_0$-VTLA 是首个在大规模触觉数据上预训练的 VTLA 模型。后训练阶段,引入可预测触觉路径,将大规模学习到的接触模式提炼为下游触觉主导操作所需的微调动作。对于离线策略优化,我们提出 ALTER,一种优势条件化离线强化学习方法,将相对进展与轨迹事件对比转化为二元优势标签,用于固定部署语料库上的策略训练,进一步提升接触密集型技能如柔性物体操作的任务特异性学习。在多个接触密集型基准测试中,$N_0$-VTLA 显著优于强基线:在全部九个真实机器人 NeoReal 任务中胜出,模拟任务二十项均值成功率达 63.8%,超过最强基线的 44.0%。使用 ALTER 训练的 $N_0$-VTLA 策略在三个长时序真实机器人任务中成功率达 75%-95%。这些结果为通用触觉驱动的操作策略奠定了基础。

原文摘要 · Abstract (English)

We present $N_0$-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, $N_0$-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, $N_0$-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. $N_0$-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.

多模态触觉感知机器人操控离线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。