arXiv:2602.18967cs.RO2026-02中稿 · ICRA

TactEx融合视觉、触觉与语言,实现类人硬度判断与可解释交互。

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

  • 融合触觉、视觉与语言,用ResNet50+LSTM处理触觉序列,跨模态对齐提升感知。
  • 在水果成熟度评估中实现90%任务成功率,各果类间分类显著差异(p<0.01)。
  • 支持自然语言指令解析与触觉输出关联的可解释性反馈,适合人机协作场景。

准确感知物体硬度对安全、灵巧的接触密集型机器人操作至关重要。本文提出TactEx,一种可解释的多模态机器人交互框架,统一视觉、触觉与语言以实现类人硬度估计与交互引导。在需触觉感知与上下文理解的水果成熟度评估任务中进行验证。系统融合GelSight-Mini触觉流、RGB观测与语言提示。采用ResNet50+LSTM模型从序列触觉数据中估计硬度,交叉模态对齐模块结合视觉线索与大语言模型(LLM)的指导。该可解释多模态界面使用户能以统计显著差异(所有果类对p<0.01)区分成熟度等级。在触觉定位方面,对比YOLO与Grounded-SAM(GSAM),发现GSAM在细粒度分割与接触点选择上更鲁棒。轻量级LLM解析用户指令,并生成与触觉输出关联的自然语言解释。端到端评估中,TactEx在简单查询下达90%任务成功率,且无需大规模调优即可泛化至新任务。结果表明,将预训练视觉与触觉模型结合语言锚定,有望推动可解释、类人的触觉感知与决策发展。

原文摘要 · Abstract (English)

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, and language for human-like hardness estimation and interactive guidance. We evaluate TactEx on fruit-ripeness assessment, a representative task that requires both tactile sensing and contextual understanding. The system fuses GelSight-Mini tactile streams with RGB observations and language prompts. A ResNet50+LSTM model estimates hardness from sequential tactile data, while a cross-modal alignment module combines visual cues with guidance from a large language model (LLM). This explainable multimodal interface allows users to distinguish ripeness levels with statistically significant class separation (p < 0.01 for all fruit pairs). For touch placement, we compare YOLO with Grounded-SAM (GSAM) and find GSAM to be more robust for fine-grained segmentation and contact-site selection. A lightweight LLM parses user instructions and produces grounded natural-language explanations linked to the tactile outputs. In end-to-end evaluations, TactEx attains 90% task success on simple user queries and generalises to novel tasks without large-scale tuning. These results highlight the promise of combining pretrained visual and tactile models with language grounding to advance explainable, human-like touch perception and decision-making in robotics.

触觉感知多模态融合可解释性人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。