聚焦角色的电影音频描述,让视障者听懂剧情关键人物与动作。
FocusedAD: Character-centric Movie Audio Description
- 通过角色感知模块追踪人物并关联姓名,实现角色中心化描述。
- 在MAD-eval-Named上达到最新性能,零样本效果也领先。
- 适合影视无障碍系统开发者,尤其关注角色叙事的场景。
电影音频描述(AD)旨在对话空白时段为盲人和视觉障碍者(BVI)讲述画面内容,相较于通用视频字幕,其要求叙述与剧情相关并明确提及角色姓名,对电影理解提出独特挑战。为识别活跃的主要角色并聚焦于情节相关区域,我们提出FocusedAD——一种角色中心的电影音频描述新框架。包含:(i) 角色感知模块(CPM),用于追踪角色区域并将其与姓名关联;(ii) 动态先验模块(DPM),通过可学习的软提示注入来自先前音频描述和字幕的上下文线索;(iii) 聚焦字幕模块(FCM),生成富含情节相关细节和命名角色的叙述。为克服角色识别局限,我们还构建了自动化角色查询库生成流程。FocusedAD在多个基准测试中取得最优表现,包括在MAD-eval-Named和我们新提出的Cinepile-AD数据集上均展现强劲零样本能力。代码与数据将公开于https://github.com/Thorin215/FocusedAD。
原文摘要 · Abstract (English)
Movie Audio Description (AD) aims to narrate visual content during dialogue-free segments, particularly benefiting blind and visually impaired (BVI) audiences. Compared with general video captioning, AD demands plot-relevant narration with explicit character name references, posing unique challenges in movie understanding.To identify active main characters and focus on storyline-relevant regions, we propose FocusedAD, a novel framework that delivers character-centric movie audio descriptions. It includes: (i) a Character Perception Module(CPM) for tracking character regions and linking them to names; (ii) a Dynamic Prior Module(DPM) that injects contextual cues from prior ADs and subtitles via learnable soft prompts; and (iii) a Focused Caption Module(FCM) that generates narrations enriched with plot-relevant details and named characters. To overcome limitations in character identification, we also introduce an automated pipeline for building character query banks. FocusedAD achieves state-of-the-art performance on multiple benchmarks, including strong zero-shot results on MAD-eval-Named and our newly proposed Cinepile-AD dataset. Code and data will be released at https://github.com/Thorin215/FocusedAD .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。