提出首个面向长视频的流式音频描述生成基准与方法,提升盲人看视频的可及性。
StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

- 将音频描述生成视为流式密集视频字幕任务,用滑动窗口实时插入描述。
- 在片段级任务上,模型达到51.0 CIDEr,比基线高10.0,全视频流式任务得分为2.4。
- 首次实现无需标注时间戳的实时生成,适合无障碍视频系统研发者使用。
视觉内容是主流传播媒介,但缺乏音频描述(ADs)时,视障人士无法访问。ADs在自然音频停顿期间叙述相关视觉事件。人工制作成本高,覆盖范围有限。现有自动方法多将任务视为视频片段字幕,依赖真实时间戳和额外上下文(如角色数据库)。当前基准也仅包含短片段与自动生成或不匹配的标注。本文提出StrAD,一个涵盖电影、纪录片、短片、演出和游戏等多样题材的长视频音频描述生成基准。将任务重新定义为流式密集视频字幕生成,通过滑动窗口处理完整视频,在无真实时间戳情况下向现有字幕流中插入描述,支持微调模型与零样本提示的视觉语言模型。在带时间戳的片段级任务中,微调模型StrAD-FT在CMD-AD上达到36.3 CIDEr(比逐帧法高10.0),在StrAD上建立51.0 CIDEr基准,在MAD-Eval上保持24.9 CIDEr竞争力。在全视频流式任务中,StrAD-FT获得SODA分数2.4,高于零样本基线的1.1,但两者在时间定位与叙事连贯性上仍存局限。此前工作以离线多阶段方式处理全视频生成,而本工作首次实现无需真实时间戳的实时生成。StrAD使全视频音频描述生成可测量,是实现可扩展无障碍的关键一步。
原文摘要 · Abstract (English)
Visual content is the dominant medium of communication, yet without audio descriptions (ADs), it remains inaccessible to blind and low-vision people. ADs narrate context-relevant visual events during natural audio pauses. Manually creating ADs is expensive, limiting coverage to a small fraction of available content. Most existing automatic AD generation methods frame the task as video clip captioning, requiring ground-truth timestamps and additional context cues such as character databases. Current benchmarks reinforce this framing, consisting of short video segments paired with automatic or task-mismatched annotations. We introduce StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games. We reformulate AD generation as streaming dense video captioning. Our approach processes full-length videos with a sliding window, inserting ADs into existing transcripts without ground-truth timestamps, and supports both fine-tuned models and zero-shot prompting of vision-language models. On the segment-level task with given timestamps, our fine-tuned StrAD-FT sets the state of the art on CMD-AD with 36.3 CIDEr (+10.0 over Shot-by-shot), establishes a reference point on StrAD (51.0 CIDEr), and remains competitive on MAD-Eval at 24.9 CIDEr. On the full-video streaming task, StrAD-FT reaches a SODA score of 2.4 against 1.1 for our zero-shot baseline StrAD-Zero, though both exhibit limitations in temporal localization and narrative coherence. While prior work has tackled full-video AD generation in an offline, multi-stage fashion, ours is the first streaming approach, generating ADs on the fly without ground-truth timestamps. StrAD makes progress on full-video AD generation measurable, a prerequisite for scaling accessibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。