用视觉数据生成车辆触觉信号,提升自动驾驶安全感知。
Synesthesia of Vehicles: Tactile Data Synthesis from Visual Inputs
- 通过跨模态时空对齐,从视觉输入预测触觉振动。
- 基于潜在扩散模型生成高质量触觉数据,时序与频域性能更优。
- 适合关注多模态融合与主动感知的自动驾驶研究者。
自动驾驶汽车依赖多模态融合保障安全,但现有视觉与光学传感器无法检测道路诱发的振动,而这对于车辆动态控制至关重要。受人类联觉启发,我们提出车辆联觉(Synesthesia of Vehicles, SoV)框架,通过视觉输入预测车辆触觉激励。我们设计了跨模态时空对齐方法以解决时间与空间差异问题,并提出基于潜在扩散的视觉-触觉联觉(VTSyn)生成模型,实现无监督高质量触觉数据合成。我们构建了一个真实车辆感知系统,在多种道路和光照条件下采集了多模态数据集。大量实验表明,VTSyn在时序、频率及分类性能上均优于现有模型,通过主动触觉感知显著提升自动驾驶安全性。
原文摘要 · Abstract (English)
Autonomous vehicles (AVs) rely on multi-modal fusion for safety, but current visual and optical sensors fail to detect road-induced excitations which are critical for vehicles' dynamic control. Inspired by human synesthesia, we propose the Synesthesia of Vehicles (SoV), a novel framework to predict tactile excitations from visual inputs for autonomous vehicles. We develop a cross-modal spatiotemporal alignment method to address temporal and spatial disparities. Furthermore, a visual-tactile synesthetic (VTSyn) generative model using latent diffusion is proposed for unsupervised high-quality tactile data synthesis. A real-vehicle perception system collected a multi-modal dataset across diverse road and lighting conditions. Extensive experiments show that VTSyn outperforms existing models in temporal, frequency, and classification performance, enhancing AV safety through proactive tactile perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。