arXiv:2510.24180cs.LG2025-10

用多模态技术自动修复视频字幕的同步与质量缺陷

V-SAT: Video Subtitle Annotation Tool

  • 融合语音、视觉与语言模型,统一检测字幕问题
  • 字幕质量提升,SUBER得分从9.6降至3.54,准确率达80%
  • 适合需要高效高质量字幕生产的媒体与平台团队

流媒体和社交媒体上音视频内容激增,对精准可访问的字幕需求随之上升。现有字幕生成方法多依赖语音转录或OCR提取,存在同步差、文本错误、格式不一、阅读速度不当及无法适应动态视听环境等问题。当前方法常仅解决单一问题,导致后期人工修正耗时费力。本文提出V-SAT(Video Subtitle Annotation Tool),一个统一框架,可自动检测并修复多种字幕质量问题。该框架结合大语言模型(LLMs)、视觉语言模型(VLMs)、图像处理与自动语音识别(ASR),利用音视频上下文线索进行判断。经验证,语言模式问题修复后SUBER得分由9.6降至3.54,图像模式问题的F1分数达约0.80。人工闭环验证保障结果质量,为鲁棒字幕标注提供了首个综合解决方案。

原文摘要 · Abstract (English)

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based extraction suffer from several shortcomings, including poor synchronization, incorrect or harmful text, inconsistent formatting, inappropriate reading speeds, and the inability to adapt to dynamic audio-visual contexts. Current approaches often address isolated issues, leaving post-editing as a labor-intensive and time-consuming process. In this paper, we introduce V-SAT (Video Subtitle Annotation Tool), a unified framework that automatically detects and corrects a wide range of subtitle quality issues. By combining Large Language Models(LLMs), Vision-Language Models (VLMs), Image Processing, and Automatic Speech Recognition (ASR), V-SAT leverages contextual cues from both audio and video. Subtitle quality improved, with the SUBER score reduced from 9.6 to 3.54 after resolving all language mode issues and F1-scores of ~0.80 for image mode issues. Human-in-the-loop validation ensures high-quality results, providing the first comprehensive solution for robust subtitle annotation.

字幕生成多模态LLM自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。