构建触觉-视觉数据集,评估柔体操作中的物理交互质量。
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

- 采集4000条专家示范,同步多视角图像、触觉信号与有限元物理状态。
- 提出变形感知成功率(DSR),要求任务完成且形变不超过阈值。
- 揭示触觉信息不自动提升性能,适合研究多模态融合与柔体操控。
物理交互质量是柔体操作的核心,但现有基准仅评估任务是否成功。策略可能完成任务却导致滑移或过度压缩。主要瓶颈在于缺乏将可见接触观测与独立物理真实状态配对的多模态数据集。本文提出SoftVTBench,一个面向物理交互感知的柔体操作视觉-触觉数据集,包含4000条专家示范和50多个资产,涵盖体积化柔体及其视觉匹配的刚体孪生体。每轮实验以20Hz频率同步多视角RGB、双指触觉RGB、标记运动轨迹、本体感知、语言指令、二进制与连续夹持动作,以及仅评估器可用的有限元(FEM)状态。基于此数据集,建立闭环基准,通过固定对象特定校准定义变形感知成功率(DSR),仅当任务完成且峰值归一化形变在容忍范围内才计为成功。在扩散策略π₀.₅与FastWAM中,12个分布内配置均有0.7%–24%的成功轨迹违反形变约束。分布外测试中,视觉-触觉模型在6组对比中任务成功率更高,在5组中DSR更高,但分布内收益混杂。结果表明,仅提供触觉信息并不保证有效多模态融合。SoftVTBench因此成为研究策略是否成功、如何物理交互及触觉何时提升交互质量的统一资源。
原文摘要 · Abstract (English)
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary bottleneck is the absence of visuo-tactile datasets that pair policy-visible contact observations with independent physical ground truth over complete tasks. We introduce SoftVTBench, a visuo-tactile dataset for physical-interaction-aware deformable-object manipulation. It contains 4,000 expert demonstrations and more than 50 assets, including volumetric deformable objects and visually matched rigid twins. At 20 Hz, each episode synchronizes multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, and binary and continuous gripper actions, alongside evaluator-only finite-element (FEM) states. Building upon this dataset, we establish a closed-loop benchmark that uses fixed object-specific calibration to define the Deformation-aware Success Rate (DSR), which counts a rollout as successful only when it completes the task and keeps peak normalized deformation within tolerance. Across Diffusion Policy, $π_{0.5}$, and FastWAM, all 12 in-distribution configurations contain successful rollouts that violate the deformation tolerance, accounting for 0.7--24% of each configuration's successes. Under distribution shift, visuo-tactile variants achieve higher task success in all six policy--suite comparisons and higher DSR in five, whereas their in-distribution benefits are mixed. These results show that making touch available does not by itself ensure effective multimodal fusion. SoftVTBench therefore provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。