arXiv:2606.11637cs.AI2026-06被引 2

构建百万级触觉数据集,让机器人理解真实世界的物理常识

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

论文配图:TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
图 1 · 摘自论文原文
  • 用多源触觉数据构建百万级推理数据集
  • 在7种传感器上覆盖415个物体,实现跨场景泛化
  • 引入动作感知表示,提升触觉信息利用效率

触觉是具身智能体理解物理世界的关键模态。尽管已有研究将触觉信号融入语言系统以进行触觉常识推理,但将其扩展到真实开放世界仍面临两大瓶颈:(1) 现有触觉推理数据集在格式和规模上有限,难以支撑从触觉观察推断物理常识,制约可迁移触觉常识的学习;(2) 触觉信号本身具有冗余性和动作依赖性,而现有方法常忽略这些特性,导致表示低效、语义表达力弱。为此,我们提出TouchThinker,从数据与表征双角度实现触觉常识推理的开放世界扩展。首先,构建了包含415个物体、8种场景、7类传感器的百万级多源触觉推理数据集TouchThinker-1M,为开放世界泛化提供坚实基础;并设计了更真实多样任务的TouchThinker-Bench开放世界评测基准。其次,提出动作感知建模机制,提升触觉表征效率,支持高效推理。实验表明,TouchThinker在多个数据集上表现优于现有先进模型。代码与数据集将开源。

原文摘要 · Abstract (English)

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering 415 objects, 8 scenarios, and 7 sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.

触觉推理多模态开放世界具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。