改进时间表示,让机器人更准预测未来动作。
Exploring Temporal Representation in Neural Processes for Multimodal Action Prediction
- 用概率生成机制融合视觉与运动信号,实现动作预测。
- 原模型在新动作上表现差,因时间表征不够鲁棒。
- 引入位置编码提升时间感知,适合长期动作预测研究。
受人类共情能力启发,本文研究条件神经过程(CNP)在机器人自监督多模态动作预测中的应用。基于镜像神经元系统(MNS)的发育研究,聚焦于自我动作预测任务。发现现有深度模态融合网络(DMBN)能通过CNP的概率生成,重建部分观测动作序列中的视听运动信号。经定性与定量评估,发现其在未见动作序列上泛化能力差,根源在于时间表示缺陷。为此提出改进版本DMBN-位置时间编码(DMBN-PTE),增强时间信息的鲁棒表示,并初步验证其扩展架构适用性的有效性。DMBN-PTE为构建可自主长期预测动作并随观测动态修正的机器人系统迈出第一步。
原文摘要 · Abstract (English)
Inspired by the human ability to understand and predict others, we study the applicability of Conditional Neural Processes (CNP) to the task of self-supervised multimodal action prediction in robotics. Following recent results regarding the ontogeny of the Mirror Neuron System (MNS), we focus on the preliminary objective of self-actions prediction. We find a good MNS-inspired model in the existing Deep Modality Blending Network (DMBN), able to reconstruct the visuo-motor sensory signal during a partially observed action sequence by leveraging the probabilistic generation of CNP. After a qualitative and quantitative evaluation, we highlight its difficulties in generalizing to unseen action sequences, and identify the cause in its inner representation of time. Therefore, we propose a revised version, termed DMBN-Positional Time Encoding (DMBN-PTE), that facilitates learning a more robust representation of temporal information, and provide preliminary results of its effectiveness in expanding the applicability of the architecture. DMBN-PTE figures as a first step in the development of robotic systems that autonomously learn to forecast actions on longer time scales refining their predictions with incoming observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。