arXiv:2505.04201cs.CV2025-05中稿 · AAAI被引 4

构建自适应触觉语言框架,提升开放场景下触觉常识推理能力

SToLa: Self-Adaptive Touch-Language Framework with Tactile Commonsense Reasoning in Open-Ended Scenarios

  • 采用混合专家架构动态融合触觉与语言模态
  • 在物理常识推理基准上表现优于现有模型
  • 适配开放场景的触觉常识推理任务

本文探讨将触觉感知融入智能系统进行多模态推理的挑战,尤其关注开放物理世界中的常识推理。我们识别出两大关键问题:模态差异性,即现有大规模触觉-语言模型常将触觉视为语言的附属子模态;以及开放触觉数据稀缺性,当前数据集缺乏多样性、开放性和复杂性以支持推理。为此,我们提出SToLa——一种自适应触觉-语言框架。SToLa利用混合专家(MoE)动态处理、统一和管理触觉与语言模态,捕捉其独特特征。尤为重要的是,我们构建了一个全面的触觉常识推理数据集与基准,包含自由形式问答、8种物理属性、4种交互特性及多样化常识知识。实验表明,SToLa在PhysiCLeAR基准及自建数据集上表现优异,验证了MoE架构在多模态管理中的有效性及其在开放场景触觉常识推理任务中的性能优势。

原文摘要 · Abstract (English)

This paper explores the challenges of integrating tactile sensing into intelligent systems for multimodal reasoning, particularly in enabling commonsense reasoning about the open-ended physical world. We identify two key challenges: modality discrepancy, where existing large touch-language models often treat touch as a mere sub-modality of language, and open-ended tactile data scarcity, where current datasets lack the diversity, open-endness and complexity needed for reasoning. To overcome these challenges, we introduce SToLa, a Self-Adaptive Touch-Language framework. SToLa utilizes Mixture of Experts (MoE) to dynamically process, unify, and manage tactile and language modalities, capturing their unique characteristics. Crucially, we also present a comprehensive tactile commonsense reasoning dataset and benchmark featuring free-form questions and responses, 8 physical properties, 4 interactive characteristics, and diverse commonsense knowledge. Experiments show SToLa exhibits competitive performance compared to existing models on the PhysiCLeAR benchmark and self-constructed datasets, proving the effectiveness of the Mixture of Experts architecture in multimodal management and the performance advantages for open-scenario tactile commonsense reasoning tasks.

触觉理解多模态推理常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。