arXiv:2604.03998cs.RO2026-04中稿 · the 2026 IEEE Inte…

让机器人实时理解语音视频指令,快速适应新命令。

VA-FastNavi-MARL: Real-Time Robot Control with Multimedia-Driven Meta-Reinforcement Learning

  • 用元强化学习将多模态指令转为可导航目标分布
  • 在多机械臂场景中样本效率显著优于基线方法
  • 无需额外计算开销,适合噪声环境下实时控制

在人机交互中,实时解析动态异构的多模态指令至关重要。我们提出VA-FastNavi-MARL框架,将异步音频-视觉输入对齐至统一潜在表示。通过将多样指令视为可导航目标的分布,利用元强化学习实现对未见指令的快速适应,推理开销几乎为零。相比依赖复杂感知处理的方法,该框架具备模态无关的流式处理能力,保障低延迟控制。在多机械臂工作空间上的验证表明,该方法在样本效率上显著优于基线,并能在噪声多模态流下保持稳定、实时执行。

原文摘要 · Abstract (English)

Interpreting dynamic, heterogeneous multimedia commands with real-time responsiveness is critical for Human-Robot Interaction. We present VA-FastNavi-MARL, a framework that aligns asynchronous audio-visual inputs into a unified latent representation. By treating diverse instructions as a distribution of navigable goals via Meta-Reinforcement Learning, our method enables rapid adaptation to unseen directives with negligible inference overhead. Unlike approaches bottlenecked by heavy sensory processing, our modality-agnostic stream ensures seamless, low-latency control. Validation on a multi-arm workspace confirms that VA-FastNavi-MARL significantly outperforms baselines in sample efficiency and maintains robust, real-time execution even under noisy multimedia streams.

机器人控制多模态元强化学习实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。