让AI通过触觉真实推理,突破视觉误导限制。
Touch-R1: Reinforcing Touch Reasoning in MLLMs

- 用触觉引导强化学习,只在真实触感下给奖励
- 在4种传感器上训练,100万组触觉数据支持
- 适合做触觉机器人、具身智能研究者
尽管基于规则的强化学习已推动多模态模型显式推理,但触觉推理仍严重不足。现有触觉-语言模型多依赖监督或对比目标,难以基于物理证据进行判断或纠正错误的视觉先验。触觉推理面临两个特有挑战:物理属性的序数性(如硬度、粗糙度)和光学触觉硬件固有的跨传感器分布偏移。本文构建了包含超过100万组同步触觉对的大规模多模态数据集TouchReason-1M,以及用于评估触觉感知与视觉-触觉冲突解决能力的严格基准TouchReason-Bench。在此基础上,提出基于Qwen2.5-VL-7B的触觉推理模型Touch-R1,采用触觉引导的GRPO目标,结合序数感知精度、跨传感器物理一致性、结构化控制及输入侧触觉接地机制。其中,触觉使用奖励仅在真实触觉输入带来更高正确率时才生效,对比去除了触觉流、打乱或噪声掩码的反事实情况。在TouchReason-Bench上,Touch-R1-7B平均优于Octopi-13B 18.4%、GPT-4o 24.7%。其结构化推理轨迹揭示了探查、比较与修正等涌现行为,证明了R1式推理可有效基于物理接触实现。
原文摘要 · Abstract (English)
While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or contrastive objectives, which limits their capacity to ground predictions in physical evidence or rectify misleading visual priors. Tactile reasoning introduces two modality-specific challenges: the ordinal nature of physical attributes (e.g., hardness, roughness) and the cross-sensor distribution shifts inherent in optical tactile hardware. In this work, we introduce TouchReason-1M, a large-scale multimodal dataset comprising over 1M synchronized tactile pairs across four distinct sensors, and TouchReason-Bench, a rigorous framework for evaluating tactile perception and visual-tactile conflict resolution. Building upon these, we propose Touch-R1, a tactile reasoning MLLM based on Qwen2.5-VL-7B. Touch-R1 is trained via a tactile-grounded GRPO objective that combines ordinal-aware accuracy, cross-sensor physical consistency, structured-format control, and an input-side tactile grounding objective. Specifically, the tactile-use reward assigns credit only when authentic tactile inputs yield superior correctness relative to counterfactual controls where the tactile stream is removed, shuffled, or noise-masked. On TouchReason-Bench, Touch-R1-7B outperforms Octopi-13B by 18.4\% and GPT-4o by 24.7\% on average. Its structured reasoning traces reveal emergent behaviors of probing, comparison, and revision, demonstrating that R1-style reasoning can be effectively grounded in physical contact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。