arXiv:2410.08860cs.CLcs.CV2024-10NAACL综述被引 12

利用大模型技术实现无障碍音频描述自动生成

Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies

  • 基于LLM和VLM的多模态生成技术实现自动音频描述
  • 可显著降低人工生成成本与时间投入
  • 适合无障碍技术、AI辅助系统研究者参考

音频描述(AD)是为视障人士在电视、电影等数字媒体场景中获取内容而设计的听觉解说。传统上由专业人员人工生成,耗时且成本高昂。近年来,自然语言处理(NLP)与计算机视觉(CV)的发展,特别是大语言模型(LLMs)和视觉-语言模型(VLMs)的突破,使自动音频描述生成成为可能。本文综述了在大模型时代下与音频描述生成相关的关键技术:探讨了先进NLP与CV技术如何应用于生成高质量音频描述,并指出了未来研究的重要方向。

原文摘要 · Abstract (English)

Audio descriptions (ADs) function as acoustic commentaries designed to assist blind persons and persons with visual impairments in accessing digital media content on television and in movies, among other settings. As an accessibility service typically provided by trained AD professionals, the generation of ADs demands significant human effort, making the process both time-consuming and costly. Recent advancements in natural language processing (NLP) and computer vision (CV), particularly in large language models (LLMs) and vision-language models (VLMs), have allowed for getting a step closer to automatic AD generation. This paper reviews the technologies pertinent to AD generation in the era of LLMs and VLMs: we discuss how state-of-the-art NLP and CV technologies can be applied to generate ADs and identify essential research directions for the future.

音频描述大模型无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。