让机器人通过预测未来声音来更好完成倒水、听音乐等操作
Learning Robot Manipulation from Audio World Models
- 用生成潜流匹配模型预测未来音频,支持长时推理
- 在倒水和听音乐任务中,有前瞻能力的系统表现更优
- 关键在于捕捉声音的节奏规律,而不仅是多模态输入
世界模型在机器人学习任务中表现出色。许多任务本质上需要多模态推理:例如往瓶子里倒水时,仅靠视觉信息可能模糊或不完整,必须结合音频随时间演变的特性,考虑其物理属性和音高模式。本文提出一种生成潜流匹配模型,用于预测未来音频观测结果,将其集成到机器人策略中后,可实现对长期后果的推理。我们在两个需感知真实环境音频或音乐信号的操作任务中验证了该系统的优越性,相比无未来展望的方法表现更佳。研究进一步表明,这些任务的成功不仅依赖多模态输入,更关键的是准确预测蕴含内在节奏模式的未来音频状态。
原文摘要 · Abstract (English)
World models have demonstrated impressive performance on robotic learning tasks. Many such tasks inherently demand multimodal reasoning; for example, filling a bottle with water will lead to visual information alone being ambiguous or incomplete, thereby requiring reasoning over the temporal evolution of audio, accounting for its underlying physical properties and pitch patterns. In this paper, we propose a generative latent flow matching model to anticipate future audio observations, enabling the system to reason about long-term consequences when integrated into a robot policy. We demonstrate the superior capabilities of our system through two manipulation tasks that require perceiving in-the-wild audio or music signals, compared to methods without future lookahead. We further emphasize that successful robot action learning for these tasks relies not merely on multi-modal input, but critically on the accurate prediction of future audio states that embody intrinsic rhythmic patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。