arXiv:2603.16086cs.ROcs.AI2026-03

让机器人实时听懂环境声音并据此执行动作,提升操作鲁棒性。

Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

  • 构建连续音频感知框架,跨执行间隙保持声音上下文
  • 在真实场景中实现90%以上声音事件的准确捕捉与响应
  • 适合需要实时声音反馈的机器人操控任务

尽管近期视觉-语言-动作(VLA)模型已开始引入音频,但通常将声音视为静态预执行提示,或仅关注人类语音。这导致在实时、以声音为中心的操作中,关键环境声学信息难以捕捉。由于低频更新或系统延迟,重要声音常被遗漏;而动作分块与开环执行更造成‘盲执行区间’,使音频事件在离散观测窗口间丢失。为此,我们提出视觉-声音-语言-动作(VSLA)连续控制范式,基于视觉、流式音频、语言和本体感知,在延迟决策循环下运行。作为实例,我们提出HEAR框架,包含四个组件:(i) 流式历史记录器,维持执行间隙间的紧凑因果音频上下文;(ii) 感知器,基于通用基础模型处理多模态输入;(iii) 推进器,以音频世界模型形式学习时间动态,预测近未来音频码;(iv) 流匹配执行器,生成平滑动作块。为解决数据稀缺问题,我们构建了用于预训练的OpenX-Sound,并推出首个以声音为中心的操控基准HEAR-Bench,具有严格的因果时序规则。实验表明,稳健的声音中心操作需依赖因果持续性和显式时序建模。该框架为具身智能体的多感官基础模型提供了实用路径,使机器人能感知并交互于动态环境。代码与视频见https://hear.irmv.top。

原文摘要 · Abstract (English)

While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. This leaves a significant gap in real-time, sound-centric manipulation where fleeting environmental acoustics provide critical state verification during task execution. Consequently, key sounds are easily missed due to low-frequency updates or system latency. This problem is exacerbated by action chunking with open-loop execution, which creates a Blind Execution Interval where acoustic events are lost between discrete audio observation windows. Recognizing the necessity of continuous auditory awareness, we formalize Vision-Sound-Language-Action (VSLA) as a continuous control paradigm conditioned on vision, streaming audio, language, and proprioception under delayed decision loops. As an instantiation, we introduce HEAR, a VSLA framework integrating four components: (i) a streaming Historizer to maintain a compact, causal audio context across execution gaps; (ii) an Envisioner adapted from omni foundation models to reason over multi-sensory inputs; (iii) an Advancer, formulated as an audio world model, to learn temporal dynamics by predicting near-future audio codes; and (iv) a flow-matching Realizer policy to generate smooth action chunks. To address the scarcity of pretraining data and evaluations for VSLA, we construct OpenX-Sound for pretraining, alongside HEAR-Bench, the first sound-centric manipulation benchmark with strict causal timing rules. Our results suggest that robust sound-centric manipulation necessitates causal persistence and explicit temporal learning. This framework provides a practical step toward multi-sensory foundation models for embodied agents, enabling robots to perceive and interact with dynamic environments. Code and videos are available at https://hear.irmv.top.

机器人操控声音感知多模态连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。