arXiv:2509.16945eess.AScs.SD2025-09

轻量化模型提升无人机语音增强,实时运行更高效。

DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement

  • 融合频带注意力与混合编码器,降低内存占用
  • 在-5至-25dB信噪比下表现优异,推理延迟低
  • 适合部署在算力受限的无人机设备上

本文提出DroFiT(Drone Frequency lightweight Transformer for speech enhancement),一种面向严重无人机自噪声环境的单麦克风语音增强网络。该模型结合频域变换器、全/子带混合编码器-解码器结构及TCN后端,实现低内存流式处理。通过可学习的跳跃门融合机制与联合频谱-时序损失函数进一步优化重建效果。模型在VoiceBank-DEMAND数据集基础上,叠加实测无人机噪声(信噪比-5至-25 dB)进行训练,并采用标准语音增强指标与计算效率评估。实验表明,DroFiT在保持竞争性增强性能的同时,显著降低计算与内存开销,为资源受限无人机平台的实时语音处理提供可行方案。音频演示样本可在官网获取。

原文摘要 · Abstract (English)

This paper proposes DroFiT (Drone Frequency lightweight Transformer for speech enhancement, a single microphone speech enhancement network for severe drone self-noise. DroFit integrates a frequency-wise Transformer with a full/sub-band hybrid encoder-decoder and a TCN back-end for memory-efficient streaming. A learnable skip-and-gate fusion with a combined spectral-temporal loss further refines reconstruction. The model is trained on VoiceBank-DEMAND mixed with recorded drone noise (-5 to -25 dB SNR) and evaluate using standard speech enhancement metrics and computational efficiency. Experimental results show that DroFiT achieves competitive enhancement performance while significantly reducing computational and memory demands, paving the way for real-time processing on resource-constrained UAV platforms. Audio demo samples are available on our demo page.

语音增强轻量化模型无人机实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。