音频生成评价难?这篇系统分析揭示了偏好学习的现状与突破方向。
Preference-Based Learning in Audio Applications: A Systematic Analysis
- 通过系统综述发现仅6%音频论文使用偏好学习,多维评估成趋势。
- 传统指标与人工判断不一致,且2021年后生成任务主导研究方向。
- 适合关注音频生成质量评估、人机偏好对齐的研究者阅读。
尽管音频与文本领域在生成模型评估上面临相似挑战,偏好学习在音频应用中仍严重不足。通过对约500篇论文的PRISMA引导系统综述发现,仅有30篇(6%)将偏好学习应用于音频任务。分析显示该领域正处于转型期:2021年前研究集中于情绪识别,采用传统排序方法(rankSVM);2021年后则转向生成任务,采用现代基于人类反馈的强化学习(RLHF)框架。我们识别出三个关键模式:(1) 多维度评估策略兴起,融合合成数据、自动指标与人工偏好;(2) 传统指标(如WER、PESQ)与人工判断在不同语境下存在不一致;(3) 多阶段训练流程逐渐成为主流,整合多种奖励信号。研究建议,尽管偏好学习在捕捉自然度、音乐性等主观品质方面具有潜力,但亟需标准化基准、更高质量数据集,并系统探究音频特有的时间因素如何影响偏好学习框架。
原文摘要 · Abstract (English)
Despite the parallel challenges that audio and text domains face in evaluating generative model outputs, preference learning remains remarkably underexplored in audio applications. Through a PRISMA-guided systematic review of approximately 500 papers, we find that only 30 (6%) apply preference learning to audio tasks. Our analysis reveals a field in transition: pre-2021 works focused on emotion recognition using traditional ranking methods (rankSVM), while post-2021 studies have pivoted toward generation tasks employing modern RLHF frameworks. We identify three critical patterns: (1) the emergence of multi-dimensional evaluation strategies combining synthetic, automated, and human preferences; (2) inconsistent alignment between traditional metrics (WER, PESQ) and human judgments across different contexts; and (3) convergence on multi-stage training pipelines that combine reward signals. Our findings suggest that while preference learning shows promise for audio, particularly in capturing subjective qualities like naturalness and musicality, the field requires standardized benchmarks, higher-quality datasets, and systematic investigation of how temporal factors unique to audio impact preference learning frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。