arXiv:2509.14659eess.AScs.LG2025-09中稿 · INTERSPEECH 2026

用人类偏好训练音频描述,让模型生成更符合人期待的文案。

Aligning Audio Captions with Human Preferences

  • 基于人类反馈强化学习,用对比语言音频模型做奖励函数。
  • 无需真实标注数据,生成描述在多数据集上更受人类青睐。
  • 适合实际应用,尤其当现有模型出错时仍能给出自然描述。

当前音频描述依赖成对的音视频-文本数据进行监督学习,但这类数据成本高且未必反映真实场景中的人类偏好。为此,我们提出一种基于人类反馈强化学习(RLHF)的偏好对齐音频描述框架。为捕捉细微偏好,我们使用人类标注的成对偏好数据训练了一个基于对比语言-音频预训练(CLAP)的奖励模型。该奖励模型被集成到强化学习框架中,用于微调任意基线描述系统,无需真实标注。跨多个数据集的大量人工评估显示,本方法生成的描述在人类评价中优于基线模型,尤其在基线模型产生错误或不自然描述时表现更佳。此外,本框架性能可媲美使用真实标注数据的监督方法,证明了与人类偏好的有效对齐及在现实场景中的可扩展性。

原文摘要 · Abstract (English)

Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To address this, we propose a preference-aligned audio captioning framework based on Reinforcement Learning from Human Feedback (RLHF). To capture nuanced preferences, we train a Contrastive Language-Audio Pretraining (CLAP) based reward model using human-labeled pairwise preference data. This reward model is integrated into an RL framework to fine-tune any baseline captioning system without ground-truth annotations. Extensive human evaluations across multiple datasets show that our method produces captions preferred over baseline models, particularly when baselines fail to provide correct and natural captions. Furthermore, our framework achieves performance comparable to supervised approaches with ground-truth data, demonstrating effective alignment with human preferences and scalability in real-world use.

音频描述强化学习人类偏好CLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。