arXiv:2507.02864cs.ROcs.CV2025-07被引 9

用生成音频让机器人在仿真中学会听声倒液,零样本迁移到真实世界。

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

  • 融合大模型与物理仿真,生成与视频同步的逼真音频。
  • 仅用仿真数据训练,实现对新容器和液体的零样本真实迁移。
  • 为难模拟的声学模态提供解决方案,适合多模态机器人研究者。

机器人需融合多种感知模态才能在真实世界有效行动,但大规模学习多模态策略仍具挑战。仿真提供可行方案,尽管视觉已受益于高保真仿真器,其他模态(如声音)却难以精确模拟。因此,目前仿真到现实的迁移主要局限于视觉任务,多模态迁移仍不成熟。本文提出MultiGen框架,将大规模生成模型集成至传统物理仿真器中,实现多感官仿真。我们在动态倒液任务上验证该框架,该任务依赖多模态反馈。通过根据仿真视频生成真实感音频,方法实现了无需真实机器人数据即可训练丰富的音视频轨迹。实验表明,该方法可实现对新容器和液体的零样本真实世界迁移,展示了生成建模在模拟难建模模态及弥合多模态仿真到现实差距方面的潜力。

原文摘要 · Abstract (English)

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from high-fidelity simulators, other modalities (e.g. sound) can be notoriously difficult to simulate. As a result, sim-to-real transfer has succeeded primarily in vision-based tasks, with multimodal transfer still largely unrealized. In this work, we tackle these challenges by introducing MultiGen, a framework that integrates large-scale generative models into traditional physics simulators, enabling multisensory simulation. We showcase our framework on the dynamic task of robot pouring, which inherently relies on multimodal feedback. By synthesizing realistic audio conditioned on simulation video, our method enables training on rich audiovisual trajectories -- without any real robot data. We demonstrate effective zero-shot transfer to real-world pouring with novel containers and liquids, highlighting the potential of generative modeling to both simulate hard-to-model modalities and close the multimodal sim-to-real gap.

多模态仿真实现生成模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。