让无人机通过视觉语言理解实时生成触觉反馈,提升人机交互沉浸感。
VLH: Vision-Language-Haptics Foundation Model
- 基于视觉语言理解生成直接触觉响应,实现感知-语言-触觉统一
- 90次实验中目标捕获成功率56.7%,纹理辨识准确率100%
- 适合虚拟现实与空中机器人交互场景,支持多模态泛化
我们提出一种新型视觉-语言-触觉基础模型VLH,将空中机器人与虚拟现实中的感知、语言和触觉反馈统一。与以往将触觉视为次要反馈通道不同,VLH将中空力与振动提示作为上下文视觉理解与自然语言指令的直接结果进行合成。系统包含一台8英寸四旋翼无人机,配备双逆五杆联动阵列用于局部触觉驱动,以及一个第一人称视角VR摄像头和第三人称俯视视角。视觉输入与语言指令由微调后的OpenVLA骨干网络处理——在自建的450个跨模态场景数据集上通过LoRA适配,输出7维动作向量(Vx, Vy, Vz, Hx, Hy, Hz, Hv)。通过INT8量化与高性能服务器,实现实时运行,频率为4-5赫兹。在人类-机器人交互实验(90次飞行)中,目标获取成功率达56.7%(平均到达时间21.3秒,位姿误差0.24米),纹理辨识准确率100%。泛化测试显示,在新任务上表现分别为:视觉70.0%、运动54.4%、物理40.0%、语义35.0%。这些结果表明,VLH能够协同演化触觉反馈与感知推理及意图表达,推动更具表现力和沉浸感的人机交互发展。
原文摘要 · Abstract (English)
We present VLH, a novel Visual-Language-Haptic Foundation Model that unifies perception, language, and tactile feedback in aerial robotics and virtual reality. Unlike prior work that treats haptics as a secondary, reactive channel, VLH synthesizes mid-air force and vibration cues as a direct consequence of contextual visual understanding and natural language commands. Our platform comprises an 8-inch quadcopter equipped with dual inverse five-bar linkage arrays for localized haptic actuation, an egocentric VR camera, and an exocentric top-down view. Visual inputs and language instructions are processed by a fine-tuned OpenVLA backbone - adapted via LoRA on a bespoke dataset of 450 multimodal scenarios - to output a 7-dimensional action vector (Vx, Vy, Vz, Hx, Hy, Hz, Hv). INT8 quantization and a high-performance server ensure real-time operation at 4-5 Hz. In human-robot interaction experiments (90 flights), VLH achieved a 56.7% success rate for target acquisition (mean reach time 21.3 s, pose error 0.24 m) and 100% accuracy in texture discrimination. Generalization tests yielded 70.0% (visual), 54.4% (motion), 40.0% (physical), and 35.0% (semantic) performance on novel tasks. These results demonstrate VLH's ability to co-evolve haptic feedback with perceptual reasoning and intent, advancing expressive, immersive human-robot interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。