从第一视角视频预测触觉信号,让机器人学会感知触摸。
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

- 基于570万张图像-触觉配对数据,从视觉直接推断触觉信息。
- 域内触力预测误差低于0.06牛,跨域接触预测优于现有方法。
- 支持零样本推理,适用于真实世界复杂场景的触觉建模。
触觉对灵巧操作至关重要,但当前广泛用于机器人学习的第一视角人类数据缺乏触觉信息。直接采集大规模触觉数据受限于传感器条件,而人类视频数据丰富、接触密集且易于扩展。这引出一个关键问题:能否仅从视觉推断触觉信号?为此,我们提出EgoTac,一种可泛化的模型,能直接从第一视角人类视频中预测丰富的触觉信息。EgoTac在超过570万组图像-触觉配对数据上训练,涵盖连续力测量与二值接触。通过学习这一多样化数据集,EgoTac捕捉了多种交互下的细微触觉动态。实验表明:域内预测平均力误差低于0.06N;在域外接触预测基准上,始终优于现有最先进接触估计器。它还能复现真实触觉数据的起伏变化,并实现对未受约束真实视频的零样本预测。扩展性分析进一步显示,数据多样性和规模提升均能稳步改善性能。总体而言,EgoTac为从第一视角人类视频中提取触觉先验提供了可扩展路径,推动通用触觉感知机器人学习。
原文摘要 · Abstract (English)
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。