arXiv:2602.15873cs.ROcs.AI2026-02被引 1

让触觉-视觉-语言模型在测试时更可靠,自动判断各模态可信度并动态调整。

Test-Time Adaptation for Tactile-Vision-Language Models

  • 根据预测不确定性和扰动响应估计每种模态的可靠性。
  • 在严重模态损坏下准确率提升最高达49.9%。
  • 适合部署于真实机器人场景中存在感知偏差的任务。

触觉-视觉-语言(TVL)模型在实际机器人和多模态感知任务中应用日益广泛,但测试时分布偏移不可避免。现有测试时自适应(TTA)方法在单模态设置下可进行过滤,却未显式处理异步跨模态偏移下的模态可靠性问题,导致部分模态不可靠时模型性能下降。本文研究此类偏移下的TVL模型TTA,提出一种可靠性感知框架,通过预测不确定性与扰动响应估计各模态可靠性,并利用该共享信号实现:(i) 过滤不可靠测试样本,(ii) 自适应融合触觉、视觉和语言特征,(iii) 以可靠性引导的目标正则化测试时优化。在TAG-C基准及额外TVL场景中,本方法持续优于强基线,严重模态损坏下准确率最高提升49.9%,凸显显式建模模态可靠性对鲁棒性测试时自适应的重要性。

原文摘要 · Abstract (English)

Tactile-vision-language (TVL) models are increasingly deployed in real-world robotic and multimodal perception tasks, where test-time distribution shifts are unavoidable. Existing test-time adaptation (TTA) methods provide filtering in unimodal settings but lack explicit treatment of modality-wise reliability under asynchronous cross-modal shifts, leaving them brittle when some modalities become unreliable. We study TTA for TVL models under such shifts and propose a reliability-aware framework that estimates per-modality reliability from prediction uncertainty and perturbation-based responses. This shared reliability signal is used to (i) filter unreliable test samples, (ii) adaptively fuse tactile, visual, and language features, and (iii) regularize test-time optimization with a reliability-guided objective. On the TAG-C benchmark and additional TVL scenarios, our approach consistently outperforms strong TTA baselines, achieving accuracy gains of up to 49.9\% under severe modality corruptions, underscoring the importance of explicit modality-wise reliability modeling for robust test-time adaptation.

多模态测试自适应机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。