让机器人理解语言、图像和地图,自动规划家务动作
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps
- 用CLIP模型统一编码语言、图像与地图信息
- 预对齐多模态嵌入空间,提升任务执行准确率
- 适合研究家庭机器人智能决策的学者参考
大型语言模型和开放词汇物体识别方法为家用服务机器人提供了更高灵活性。通过提供任务描述和环境信息,无需单独实现每项任务即可应对家庭任务的多样性。本文提出LIAM——一种端到端模型,根据语言、图像、动作和语义地图输入预测动作序列。语言与图像输入使用CLIP主干网络编码,并设计两种预训练任务以微调权重并预对齐潜在空间。在ALFRED数据集(模拟生成的家庭任务基准)上评估表明,多模态嵌入空间的预对齐至关重要,且引入语义地图显著提升了性能。
原文摘要 · Abstract (English)
The availability of large language models and open-vocabulary object perception methods enables more flexibility for domestic service robots. The large variability of domestic tasks can be addressed without implementing each task individually by providing the robot with a task description along with appropriate environment information. In this work, we propose LIAM - an end-to-end model that predicts action transcripts based on language, image, action, and map inputs. Language and image inputs are encoded with a CLIP backbone, for which we designed two pre-training tasks to fine-tune its weights and pre-align the latent spaces. We evaluate our method on the ALFRED dataset, a simulator-generated benchmark for domestic tasks. Our results demonstrate the importance of pre-aligning embedding spaces from different modalities and the efficacy of incorporating semantic maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。