让机器人通过声音主动识别并操作物体,无需文字指令。
From Instruction to Event: Sound-Triggered Mobile Manipulation
- 用声音触发而非文字指令驱动机器人操作
- 在双声源干扰下仍能准确识别主声源并完成交互
- 适合需要自主感知的移动操作场景,如家庭服务机器人
当前移动操作研究多依赖预设文本指令,使机器人处于被动状态,限制其自主性与环境响应能力。为此,本文提出声触发式移动操作,使机器人无需显式指令即可主动感知并响应发声物体。为支持该任务,构建了融合声学渲染与物理交互的Habitat-Echo数据平台,并设计了一套包含高层任务规划器与底层策略模型的基线方法。大量实验表明,该基线使机器人能够主动探测并响应听觉事件,无需逐项指令。尤其在双声源挑战场景中,机器人成功从重叠声干扰中分离出主声源并完成首次交互,随后继续操作次级物体,验证了方法的鲁棒性。
原文摘要 · Abstract (English)
Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to execute tasks. However, this setting confines agents to a passive role, limiting their autonomy and ability to react to dynamic environmental events. To address these limitations, we introduce sound-triggered mobile manipulation, where agents must actively perceive and interact with sound-emitting objects without explicit action instructions. To support these tasks, we develop Habitat-Echo, a data platform that integrates acoustic rendering with physical interaction. We further propose a baseline comprising a high-level task planner and low-level policy models to complete these tasks. Extensive experiments show that the proposed baseline empowers agents to actively detect and respond to auditory events, eliminating the need for case-by-case instructions. Notably, in the challenging dual-source scenario, the agent successfully isolates the primary source from overlapping acoustic interference to execute the first interaction, and subsequently proceeds to manipulate the secondary object, verifying the robustness of the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。