让视频旁白连贯流畅,避免重复描述同一画面
More than a Moment: Towards Coherent Sequences of Audio Descriptions
- 生成多个候选描述后逐句筛选,确保前后连贯
- 新指标StoryRecall衡量整体叙事还原度,优于以往方法
- 无需训练,适合无障碍视频生成场景
音频描述(AD)为视障人群提供视频关键信息,但现有自动方法独立生成每段描述,导致内容重复、缺乏连贯性。为此,我们提出无需训练的CoherentAD方法:先对每个时间区间生成多个候选描述,再通过自回归方式选择序列,形成连贯且信息丰富的叙事。为全面评估,引入序列级指标StoryRecall,衡量预测描述对真实叙事的还原能力,并结合重复率指标捕捉连续输出中的冗余问题。实验表明,该方法显著提升叙事连贯性与理解度,优于依赖独立生成的基线模型。
原文摘要 · Abstract (English)
Audio Descriptions (ADs) convey essential on-screen information, allowing visually impaired audiences to follow videos. To be effective, ADs must form a coherent sequence that helps listeners to visualise the unfolding scene, rather than describing isolated moments. However, most automatic methods generate each AD independently, often resulting in repetitive, incoherent descriptions. To address this, we propose a training-free method, CoherentAD, that first generates multiple candidate descriptions for each AD time interval, and then performs auto-regressive selection across the sequence to form a coherent and informative narrative. To evaluate AD sequences holistically, we introduce a sequence-level metric, StoryRecall, which measures how well the predicted ADs convey the ground truth narrative, alongside repetition metrics that capture the redundancy across consecutive AD outputs. Our method produces coherent AD sequences with enhanced narrative understanding, outperforming prior approaches that rely on independent generations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。