构建音视频角色库,实现动画片角色自动识别与无障碍观看。
Character-Centric Understanding of Animated Movies
- 用网络数据自动构建包含音视频的角色样本库,支持多模态识别。
- 在75部动画片上验证,显著提升视障者语音描述与听障者字幕的准确性。
- 适合做动画内容无障碍化、角色追踪的研究人员使用。
动画电影因其独特的角色设计和想象力叙事而引人入胜,但对现有识别系统构成挑战。与传统人脸识别中一致的视觉模式不同,动画角色在外观、动作和形变上具有极强多样性。本文提出一种音视频联动管道,实现动画角色的自动且鲁棒的识别,从而增强对动画片的角色中心理解。核心是利用在线资源自动构建音视频角色库,包含每个角色的视觉样例和语音样本,支持在长尾分布下进行多模态识别。基于准确的角色识别,我们探索了两个下游应用:为视障人群生成语音描述(AD),以及为听障人群提供角色感知字幕。为推动该领域研究,我们引入了CMD-AM数据集,涵盖75部动画电影的全面标注。相比基于人脸检测的方法,本方法在可访问性和叙事理解方面均有显著提升。代码与数据集详见https://www.robots.ox.ac.uk/~vgg/research/animated_ad/。
原文摘要 · Abstract (English)
Animated movies are captivating for their unique character designs and imaginative storytelling, yet they pose significant challenges for existing recognition systems. Unlike the consistent visual patterns detected by conventional face recognition methods, animated characters exhibit extreme diversity in their appearance, motion, and deformation. In this work, we propose an audio-visual pipeline to enable automatic and robust animated character recognition, and thereby enhance character-centric understanding of animated movies. Central to our approach is the automatic construction of an audio-visual character bank from online sources. This bank contains both visual exemplars and voice (audio) samples for each character, enabling subsequent multi-modal character recognition despite long-tailed appearance distributions. Building on accurate character recognition, we explore two downstream applications: Audio Description (AD) generation for visually impaired audiences, and character-aware subtitling for the hearing impaired. To support research in this domain, we introduce CMD-AM, a new dataset of 75 animated movies with comprehensive annotations. Our character-centric pipeline demonstrates significant improvements in both accessibility and narrative comprehension for animated content over prior face-detection-based approaches. For the code and dataset, visit https://www.robots.ox.ac.uk/~vgg/research/animated_ad/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。