arXiv:2412.00342cs.AI2024-12被引 2

用大模型提升聋哑人群视频字幕准确率,效果显著。

Empowering the Deaf and Hard of Hearing Community: Enhancing Video Captions Using Large Language Models

  • 用大模型修正语音识别生成的字幕,提升语义准确性。
  • 改进后字幕词错误率降至9.75%,比原字幕降低57.72%。
  • 特别适合改善聋哑人群观看视频时的可及性体验。

在数字时代,视频内容成为信息、教育与娱乐的主要来源。然而,听障及重听(DHH)群体常因自动语音识别(ASR)系统生成的字幕不准确而难以获取视频信息。本文提出一种基于大语言模型(LLM)的字幕增强方法,旨在提升字幕的准确性和上下文理解能力。我们构建了一种新流程,利用GPT-3.5和Llama2-13B等先进大模型对ASR输出进行纠错。为评估实际应用效果,我们建立了一个反映真实使用场景的数据集。实验表明,经大模型优化后的字幕显著提升质量:ChatGPT-3.5的词错误率(WER)降至9.75%,相比原始ASR字幕(WER: 23.07%)降低约57.72%。

原文摘要 · Abstract (English)

In today's digital age, video content is prevalent, serving as a primary source of information, education, and entertainment. However, the Deaf and Hard of Hearing (DHH) community often faces significant challenges in accessing video content due to the inadequacy of automatic speech recognition (ASR) systems in providing accurate and reliable captions. This paper addresses the urgent need to improve video caption quality by leveraging Large Language Models (LLMs). We present a comprehensive study that explores the integration of LLMs to enhance the accuracy and context-awareness of captions generated by ASR systems. Our methodology involves a novel pipeline that corrects ASR-generated captions using advanced LLMs. It explicitly focuses on models like GPT-3.5 and Llama2-13B due to their robust performance in language comprehension and generation tasks. We introduce a dataset representative of real-world challenges the DHH community faces to evaluate our proposed pipeline. Our results indicate that LLM-enhanced captions significantly improve accuracy, as evidenced by a notably lower Word Error Rate (WER) achieved by ChatGPT-3.5 (WER: 9.75%) compared to the original ASR captions (WER: 23.07%), ChatGPT-3.5 shows an approximate 57.72% improvement in WER compared to the original ASR captions.

字幕生成大模型听障辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。