arXiv:2409.10048cs.SDcs.AI2024-09中稿 · ICASSP 2025被引 3

用音频驱动强化学习,让机器人自动转向说话人。

Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments

  • 基于立体语音信号,用深度Q学习实现自动转向
  • 无混响环境下表现接近完美,有混响时仍优于随机行为
  • 训练环境混响程度影响泛化能力,高混响训练更鲁棒

尽管深度强化学习(DRL)在音频信号处理中取得显著进展,但在人机交互场景下,如导航、视线控制和头部朝向控制等任务中,音频驱动的DRL研究仍较少。本文提出一种音频驱动的DRL框架,利用深度Q学习,使智能体根据立体语音记录自主朝向说话人。实验表明,该智能体在无混响环境(anechoic)下训练时可近乎完美地完成任务;而在自然声学环境中存在混响时,性能下降,但仍显著优于随机行为基线。进一步量化发现:在中或高混响环境下训练的策略能泛化到低混响环境,但反之则不可行。结果表明,音频驱动DRL在头部朝向控制中具有潜力,且需设计更鲁棒的训练策略以实现跨环境泛化。

原文摘要 · Abstract (English)

Although deep reinforcement learning (DRL) approaches in audio signal processing have seen substantial progress in recent years, audio-driven DRL for tasks such as navigation, gaze control and head-orientation control in the context of human-robot interaction have received little attention. Here, we propose an audio-driven DRL framework in which we utilise deep Q-learning to develop an autonomous agent that orients towards a talker in the acoustic environment based on stereo speech recordings. Our results show that the agent learned to perform the task at a near perfect level when trained on speech segments in anechoic environments (that is, without reverberation). The presence of reverberation in naturalistic acoustic environments affected the agent's performance, although the agent still substantially outperformed a baseline, randomly acting agent. Finally, we quantified the degree of generalization of the proposed DRL approach across naturalistic acoustic environments. Our experiments revealed that policies learned by agents trained on medium or high reverb environments generalized to low reverb environments, but policies learned by agents trained on anechoic or low reverb environments did not generalize to medium or high reverb environments. Taken together, this study demonstrates the potential of audio-driven DRL for tasks such as head-orientation control and highlights the need for training strategies that enable robust generalization across environments for real-world audio-driven DRL applications.

强化学习音频驱动机器人交互泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。