用大模型生成语音质量评分,让音频视觉增强更符合人耳感受。
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
- 用语音大模型生成描述性反馈,转化为可优化的奖励信号。
- 在AVSEC-4数据集上,各项指标均优于监督和基于DNSMOS的强化学习方法。
- 适合关注语音感知质量、追求可解释优化的研究者。
现有音频-视觉语音增强(AVSE)方法多采用尺度不变信噪比(SI-SNR)或均方误差(MSE)作为目标,但这些指标与主观语音质量的相关性较差,且优化过程缺乏可解释性。本文提出一种基于强化学习的AVSE框架,引入大语言模型(LLM)构建可解释的奖励模型:由音频语言模型生成增强后语音的自然语言描述,再通过情感分析模型转化为1-5分的评分,作为PPO算法的奖励信号,用于微调预训练的AVSE模型。相比传统标量指标,该方法生成的反馈语义丰富,能明确描述语音质量提升。在AVSEC-4数据集上的实验表明,该方法在PESQ、STOI、神经质量评估指标及主观听感测试中均优于监督基线和基于DNSMOS的强化学习基线。
原文摘要 · Abstract (English)
In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, their correlation with perceived speech quality is often suboptimal and provides limited interpretability for optimization. This work proposes a reinforcement learning-based AVSE framework with a Large Language Model (LLM)-based interpretable reward model. An audio LLM generates natural language descriptions of enhanced speech, which are converted by a sentiment analysis model into a 1-5 rating score serving as the PPO reward for fine-tuning a pretrained AVSE model. Compared with scalar metrics, LLM-generated feedback is semantically rich and explicitly describes speech quality improvements. Experiments on the AVSEC-4 dataset show that the proposed method outperforms a supervised baseline and a DNSMOS-based RL baseline in PESQ, STOI, neural quality metrics, and subjective listening tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。