arXiv:2602.11934cs.RO2026-02

让机器人精准抓取,靠的是能感知接触细节的视觉特征。

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation

  • 用确定性方法提取扩散模型的对应关系特征
  • 在真实机器人上完成接触敏感任务,成功率显著提升
  • 适合需要精细触控的机器人操作场景

机器人操作常在最后几毫米失败:策略虽识别出目标物体,却忽略姿态偏差、边界或接触前对齐。问题源于语义不变性掩盖了闭环控制所需的对应线索,或这些线索未以可用形式传递给策略。现代视觉编码器提供强语义抽象,但接触密集型操作需具备对应敏感性:对姿态、边界和接触几何变化具有区分响应的特征。扩散特征具备强大稠密对应先验,但因随机性、延迟和表征漂移难以直接使用。我们提出Robot-DIFT,一种用于实时控制的确定性扩散衍生主干网络。通过流形蒸馏,将噪声条件扩散教师模型转化为纯净输入、单次通过的学生模型,同时保留教师特征流形。空间-语义特征金字塔网络(S2-FPN)融合粗到细的学生解码器特征,生成暴露语义上下文与精细接触细节的视觉令牌。在RoboCasa、LIBERO-10及真实机器人上,Robot-DIFT在接触敏感任务中超越视觉-语言、自监督、几何导向及扩散基线。受控主干/读出替换实验表明,S2-FPN解锁而非取代扩散对应先验。

原文摘要 · Abstract (English)

Robot manipulation often fails in the final millimeters: a policy may recognize the right object yet miss the pose offsets, boundaries, or pre-contact alignments needed for action. We argue that such failures arise when semantic invariance suppresses correspondence cues for closed-loop control, or when these cues are not exposed to the policy in a usable form. Modern visual encoders provide strong semantic abstractions, but contact-rich manipulation requires correspondence sensitivity: discriminative feature responses to action-relevant changes in pose, boundary, and contact geometry. Diffusion features provide a strong prior for dense correspondence, but direct use is impractical due to stochasticity, latency, and representation drift. We introduce Robot-DIFT, a deterministic diffusion-derived backbone for real-time control. Through Manifold Distillation, Robot-DIFT converts a noise-conditioned diffusion Teacher into a clean-input, single-pass Student while preserving the teacher's feature manifold. A Spatial--Semantic Feature Pyramid Network (S2-FPN) fuses coarse-to-fine Student decoder features into visual tokens that expose semantic context and fine contact detail to the policy. Across RoboCasa, LIBERO-10, and real robots, Robot-DIFT outperforms vision--language, self-supervised, geometry-oriented, and diffusion baselines on contact-sensitive tasks. Controlled backbone/readout swaps show that S2-FPN unlocks, rather than replaces, the diffusion correspondence prior.

机器人操作视觉特征扩散模型接触感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。