arXiv:2606.13970cs.ROcs.LG2026-06

提出可处理缺失模态的注意力模型,提升机器人在不完备数据下的预测鲁棒性。

An Attention-based Model for Robust Forecasting with Missing Modality

论文配图:An Attention-based Model for Robust Forecasting with Missing Modality
图 1 · 摘自论文原文
  • 基于条件变分自编码器与Transformer,用注意力机制融合多模态数据
  • 在训练和推理中均支持缺失模态,仍能学习完整统一表征
  • 在5个数据集上验证,优于现有多模态融合方法,适合真实机器人应用

多模态机器人学习中的缺失模态问题是核心挑战,因现实机器人系统常面临传感器数据不全。尽管基于注意力的模型因可使用单一主干网络处理多模态数据而备受青睐,但多数模型假设训练和推理时所有模态均可用,限制了其在机器人感知与决策中的应用。本文提出一种新型多模态模型,可在训练和推理阶段处理缺失模态。该模型采用条件变分自编码器(CVAE)形式,结合Transformer架构,利用注意力机制学习统一的固定维度表示,即使部分模态缺失也能有效工作。实验表明,该模型能在缺失模态条件下训练并逼近所有模态的稳健表示。我们在两个机器人学习任务(人类轨迹预测与机器人操作预测)的五个多模态数据集上评估该方法,结果证明其能有效从不完整数据中学习,显著优于现有融合方法。

原文摘要 · Abstract (English)

Learning with missing modalities is a fundamental challenge in multimodal robot learning, as real-world robotic systems often operate in environments with incomplete sensor data. Attention-based models are appealing for processing multimodal data because they can handle multiple modalities with a single backbone network. However, most multimodal models assume that all modalities are available during both training and inference, limiting their applicability in robotic perception and decision-making. In this paper, we introduce a multimodal model designed to handle missing modalities during both training and inference. The model is formulated as a conditional variational autoencoder (CVAE) and incorporates a transformer-based architecture that leverages attention mechanisms to learn a unified, fixed-dimensional representation, even when some modalities are missing. We show that our proposed model can be trained with missing modalities while approximating a robust representation of all modalities. We evaluate our approach on five multimodal datasets across two robot learning tasks: human trajectory prediction and robot manipulation forecasting. Experimental results demonstrate that our model effectively learns from incomplete data and is superior to prior multimodal fusion approaches.

多模态注意力机制机器人学习缺失数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。