让音频描述连贯讲好故事,无需训练就能记住角色和情节
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

- 用可验证的叙事记忆跨场景保持角色与事件关联
- 在多个基准上提升故事连贯性与事实准确率
- 适合无障碍影视辅助,无需剧本或额外标注
长时序音频描述不仅需描述可见动作,还需持续保留角色、事件、关系与叙事背景,使视障和低视力观众能理解影片。现有视频-语言模型对短片段有效,但常孤立处理每一时刻,导致描述缺乏角色身份、事件意义及前后关联。本文提出StoryTeller,一种无需训练的叙事感知长时序音频描述框架。它不依赖局部视觉线索,而是维护一个经语义过滤与VLM验证的可信叙事记忆,将关键信息跨场景传递,确保后续描述保持连贯、真实且具上下文信息。仅需原始视频与片名,可选获取公开电影元数据以确认人物与背景,但只采纳视频支持的事实。该方法无需字幕、脚本、描述文本、对齐字幕、角色库、预计算人脸身份或任务微调。为评估生成描述是否保留叙事信息,我们引入StoryAD-QA——一个基于问答的评测基准,测试语言模型能否仅凭生成描述回答叙事问题。在标准音频描述基准和多样长视频上的实验表明,StoryTeller在自动评估、问答评估和人工评估中,均显著优于强基线模型,提升叙事连贯性、事实准确性和故事理解能力。
原文摘要 · Abstract (English)
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。