用形状和接触姿态控制,生成多模态触觉图像。
MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose
- 用深度图和结构化提示双重条件,统一生成多种触觉传感器图像。
- 在8个物体上,性能比基线提升超60%,最大达134.6%。
- 合成数据可大幅减少真实数据需求,适合机器人触觉研究。
获取对齐的视觉-触觉数据集耗时且成本高,需专用硬件与大规模采集。合成生成具潜力,但现有方法多为单模态,限制跨模态学习。我们提出MultiDiffSense,一个统一扩散模型,可在单一架构内生成多种视觉触觉传感器(ViTac、TacTip、ViTacTip)的图像。方法基于CAD生成的对齐深度图与编码传感器类型及4自由度接触姿态的结构化提示进行双重条件控制,实现可控且物理一致的多模态合成。在8个物体(5个已见,3个新物体)及未见姿态上评估,相比Pix2Pix cGAN基线,SSIM提升分别为+36.3%(ViTac)、+134.6%(ViTacTip)和+64.7%(TacTip)。下游3-DoF姿态估计任务中,混合50%合成与50%真实数据,可将所需真实数据量减半,同时保持竞争力。MultiDiffSense缓解了触觉感知中的数据收集瓶颈,支持可扩展、可控的多模态数据集生成,适用于机器人应用。
原文摘要 · Abstract (English)
Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning. We present MultiDiffSense, a unified diffusion model that synthesises images for multiple vision-based tactile sensors (ViTac, TacTip, ViTacTip) within a single architecture. Our approach uses dual conditioning on CAD-derived, pose-aligned depth maps and structured prompts that encode sensor type and 4-DoF contact pose, enabling controllable, physically consistent multi-modal synthesis. Evaluating on 8 objects (5 seen, 3 novel) and unseen poses, MultiDiffSense outperforms a Pix2Pix cGAN baseline in SSIM by +36.3% (ViTac), +134.6% (ViTacTip), and +64.7% (TacTip). For downstream 3-DoF pose estimation, mixing 50% synthetic with 50% real halves the required real data while maintaining competitive performance. MultiDiffSense alleviates the data-collection bottleneck in tactile sensing and enables scalable, controllable multi-modal dataset generation for robotic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。