用非线性声学+强化学习提升机器人在嘈杂环境下的听觉交互能力
A Synergistic Framework of Nonlinear Acoustic Computing and Reinforcement Learning for Real-World Human-Robot Interaction
- 结合物理声学方程与强化学习,动态优化声场参数
- 实测显示噪声抑制更强、延迟更低、多语种识别更准
- 适合做智能机器人、可穿戴设备等实时听觉系统
本文提出一种融合非线性声学计算与强化学习的新框架,用于复杂噪声和混响条件下的真实人机交互。利用物理驱动的波动方程(如Westervelt、KZK模型),系统可捕捉谐波生成、激波形成等高阶声学现象。通过将这些模型嵌入强化学习控制环路,系统自适应优化吸收率、波束成形等关键参数,有效缓解多路径干扰与非平稳噪声。实验涵盖远场定位、弱信号检测和多语言语音识别,在严苛真实场景中表现优异,显著优于传统线性方法和纯数据驱动基线,实现更强噪声抑制、极低延迟与鲁棒准确率。该系统在人工智能硬件、机器人、机器听觉、人工听觉及脑机接口等领域具有广泛应用前景。
原文摘要 · Abstract (English)
This paper introduces a novel framework integrating nonlinear acoustic computing and reinforcement learning to enhance advanced human-robot interaction under complex noise and reverberation. Leveraging physically informed wave equations (e.g., Westervelt, KZK), the approach captures higher-order phenomena such as harmonic generation and shock formation. By embedding these models in a reinforcement learning-driven control loop, the system adaptively optimizes key parameters (e.g., absorption, beamforming) to mitigate multipath interference and non-stationary noise. Experimental evaluations, covering far-field localization, weak signal detection, and multilingual speech recognition, demonstrate that this hybrid strategy surpasses traditional linear methods and purely data-driven baselines, achieving superior noise suppression, minimal latency, and robust accuracy in demanding real-world scenarios. The proposed system demonstrates broad application prospects in AI hardware, robot, machine audition, artificial audition, and brain-machine interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。