arXiv:2608.09765cs.CLcs.CV2026-08被引 1

提出真实电影音频描述新任务,让模型自定何时何地描述画面。

REFRAMED: Towards Realistic Audio Description Generation for Movies

  • 让模型自主决定描述内容与时机,而非预设时间点。
  • 构建含3302场景的高质量数据集,含英美双版本专业配音。
  • 提供多参考对比评测,推动视频理解研究进展。

音频描述(AD)是为视障观众提供的视频关键视觉内容口头解说。与标准视频字幕不同,AD是一项结构化编辑任务:描述必须插入对话语境间隙,并仅传达理解叙事所必需的信息。现有方法将AD生成置于人为设定的内容与时间条件下,使其退化为片段级字幕生成,且依赖噪声较大的语音转录与对齐流程,缺乏建模叙事上下文所需的丰富平行数据。本文提出一种新范式:模型需联合决策描述内容与时机。为此,我们构建了REFRAMED数据集,包含206部电影中的2,023个视频、3,302个场景,配有专业级美国与英国双版本音频描述、专业字幕及对齐剧本。同时提供人工标注的挑战集,涵盖完整影片与多份参考描述,并设计基于对话空隙与多参考比较的评估协议。实验表明,当前先进系统与多模态大模型虽优于基线,但仍远低于人类专家水平。该数据集与基准为视频理解研究奠定新基础。

原文摘要 · Abstract (English)

Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,023 videos that span 3,302 scenes from 206 movies, with professional AD transcripts (both American and British versions), professional subtitles and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that they outperform trivial baselines but fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on video understanding.

音频描述视频理解多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。