arXiv:2602.13833cs.RO2026-02中稿 · RSS 2026

用3D触觉场实现工具操作的跨类别泛化,兼顾语义与物理接触信息。

Semantic-Contact Fields for Category-Level Generalizable Tactile Tool Manipulation

  • 构建融合视觉语义与密集接触概率/力的3D统一表示
  • 在仿真预训练+少量真实数据微调后,对未见过的工具泛化性能显著提升
  • 适合需要高精度触觉控制的机器人操作任务研究者

通用工具操作需兼具语义规划与精确物理控制。当前通用机器人策略如视觉-语言-动作(VLA)模型缺乏物理接地能力,而依赖触觉感知的现有方法通常仅适用于特定实例,难以跨不同工具几何形状泛化。为弥合此差距,我们提出语义-接触场(SCFields),一种融合视觉语义与密集外在接触估计(包括接触概率和力)的统一3D表示。通过两阶段仿真到现实接触学习流程:先在大规模仿真中预训练以学习几何感知的接触先验,再利用少量真实数据(通过几何启发式与力优化伪标注)微调,使真实触觉信号对齐。所得力感知表示作为扩散策略的密集观测输入,实现对未见工具实例的物理泛化。在刮擦、蜡笔绘画和剥皮任务上的实验表明,该方法显著优于纯视觉与原始触觉基线,展现出鲁棒的类别级泛化能力。

原文摘要 · Abstract (English)

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often lack the physical grounding required for contact-rich tool manipulation. Conversely, existing contact-aware policies that leverage tactile or haptic sensing are typically instance-specific and fail to generalize across diverse tool geometries. Bridging this gap requires learning representations that are both semantically transferable and physically grounded, yet a fundamental barrier remains: diverse real-world tactile data are prohibitive to collect at scale, while direct zero-shot sim-to-real transfer is challenging due to the complex nonlinear deformation of soft tactile sensors. To address this, we propose Semantic-Contact Fields (SCFields), a unified 3D representation that fuses visual semantics with dense extrinsic contact estimates, including contact probability and force. SCFields is learned through a two-stage Sim-to-Real Contact Learning Pipeline: we first pre-train on large-scale simulation to learn geometry-aware contact priors, then fine-tune on a small set of real data pseudo-labeled via geometric heuristics and force optimization to align real tactile signals. The resulting force-aware representation serves as the dense observation input to a diffusion policy, enabling physical generalization to unseen tool instances. Experiments on scraping, crayon drawing, and peeling demonstrate robust category-level generalization, significantly outperforming vision-only and raw-tactile baselines. Project page: https://kevinskwk.github.io/SCFields/.

触觉控制机器人操作泛化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。