arXiv:2606.22766cs.CV2026-06

用强化学习提升音频描述的准确与连贯性,效果显著优于现有方法。

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

论文配图:READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
图 1 · 摘自论文原文
  • 采用强化学习框架,从序列层面优化生成内容
  • 在多个数据集上超越现有方法,显著提升准确率与连贯性
  • 适合关注无障碍多媒体生成的研究者与开发者

音频描述旨在为视障及低视力人群生成简洁的视觉内容解说。现有方法或依赖预训练多模态模型,风格不匹配;或基于逐词预测训练,未能充分挖掘模型潜力且易生成通用表达。本文提出READ,首个基于强化学习的音频描述生成框架。READ将生成任务建模为序列级优化,引入参考匹配、长度和格式奖励,并设计上下文感知的连贯性奖励以提升叙述一致性。在MAD-Eval、CMD-AD和TV-AD三个数据集上的实验表明,READ在多种评估指标上均显著优于先前方法。结果验证了强化学习在生成准确、连贯音频描述方面的潜力。代码、模型与评测结果将公开发布。

原文摘要 · Abstract (English)

Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on prompting off-the-shelf multimodal models, which often mismatch AD style, or partially optimize training-based systems with next-token prediction, which under-explores model capacity and biases generation toward generic expressions. We present READ, the first reinforcement-learning (RL) framework for training-based AD generation. READ formulates AD as sequence-level optimization with reference-matching, length, and format rewards, and further introduces a dedicated coherence reward under context-aware supervision to promote narratively coherent descriptions. Experiments on MAD-Eval, CMD-AD, and TV-AD show that READ substantially outperforms prior methods across diverse evaluation metrics. Our results highlight RL as a promising paradigm for accurate and coherent AD generation. Our codes, models, and benchmark results will be publicly available.

音频描述强化学习多模态无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。