arXiv:2609.07440cs.RO2026-09

无需标注数据,自动分离机器人行走噪音,保留环境声音

Open-Set Ego-Noise Separation for Legged-Robot Audition via Annotation-Free Adaptation and Pretrained-Model Transfer

论文配图:Open-Set Ego-Noise Separation for Legged-Robot Audition via Annotation-Free Adaptation and Pretrained-Model Transfer
图 1 · 摘自论文原文
  • 用无标注数据自动选中含噪音的音频片段
  • 混合环境音生成训练对,提升分离精度
  • 适合没条件收集噪音数据的机器人听觉系统

本文提出一种开集机器人自身噪声分离框架,通过无标注自适应与预训练模型迁移实现。该框架在不依赖预先指定环境声音类别的情况下,去除机器人自身产生的行走噪声(如脚步撞击、关节松动、电机噪声),同时保留环境声音。声学感知可提供视觉之外的环境线索,但行走引发的自身噪声严重干扰录音。框架首先利用RecurGraph方法,通过聚合音频片段嵌入并构建嵌入图,传播得分,从无标签录音中筛选出以自身噪声为主导的片段。这些片段与大规模声事件数据集中的多样化环境声音混合,生成成对的混合-目标监督信号,用于开集分离训练。随后,Transfer-DiT将通用零样本神经分离器适配到目标机器人,实现高保真开集自身噪声分离。在双足与四足机器人上的实验表明,该方法能可靠选择有效片段,显著提升分离质量及下游任务性能。结果验证了无需单独录制自身噪声数据或人工片段标注即可实现自适应的可行性。

原文摘要 · Abstract (English)

This paper proposes an open-set ego-noise separation framework for legged-robot audition via annotation-free adaptation and pretrained-model transfer. The framework removes robot-specific ego-noise while preserving environmental sounds whose classes are not specified in advance. Acoustic sensing provides cues about a robot's surroundings beyond the visual field, but walking-induced ego-noise from footstep impacts, joint-backlash rattling, and motor noise severely contaminates the recordings. The framework first uses RecurGraph to select ego-noise-dominant clips from the unlabeled recordings by aggregating clip embeddings into an embedding centroid and propagating scores over an audio-embedding graph. The selected clips are mixed with diverse environmental sounds from a large-scale sound-event dataset to provide paired mixture--target supervision for open-set separation. Transfer-DiT then adapts a general-purpose zero-shot neural separator to achieve high-fidelity open-set ego-noise separation for the target robot. Experiments with bipedal and quadrupedal robots show reliable clip selection and improvements in separation quality and downstream task performance over baseline separators. These results demonstrate the feasibility of annotation-free adaptation without separately recorded ego-noise-only data or manual clip-level annotations.

语音分离机器人听觉无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。