arXiv:2512.19161cs.CL2025-12中稿 · KDD被引 1

评测主流语音识别模型在意大利电视字幕中的表现,发现需人工介入才够用。

From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs

  • 在50小时意大利语电视片段上对比4个主流语音识别模型
  • 模型错误率仍高于专业字幕员,无法完全替代人工
  • 提出云端人机协同系统,提升字幕制作效率

字幕对视频可访问性和观众参与度至关重要。当前基于编码器-解码器架构的自动语音识别(ASR)系统在标准基准数据集上已显著降低转录错误,但在真实生产环境,尤其是长时长意大利语视频等非英语内容中,其表现仍缺乏研究。本文针对一家意大利媒体公司,开展了一个专业字幕系统构建的案例研究。为指导系统设计,我们在50小时意大利电视节目数据集上评估了四个前沿ASR模型(Whisper Large v2、AssemblyAI Universal、Parakeet TDT v3 0.6b和WhisperX),并与专业人工字幕员的表现进行对比。结果表明,尽管当前模型尚无法满足媒体行业对全自动字幕的准确性要求,但可作为显著提升人工效率的有效工具。研究结论强调人机协同(HITL)模式的关键作用,并展示了我们设计的生产级云端基础设施以支持该工作流。

原文摘要 · Abstract (English)

Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively reduced transcription errors on standard benchmark datasets. However, their performance in real-world production environments, particularly for non-English content like long-form Italian videos, remains largely unexplored. This paper presents a case study on developing a professional subtitling system for an Italian media company. To inform our system design, we evaluated four state-of-the-art ASR models (Whisper Large v2, AssemblyAI Universal, Parakeet TDT v3 0.6b, and WhisperX) on a 50-hour dataset of Italian television programs. The study highlights their strengths and limitations, benchmarking their performance against the work of professional human subtitlers. The findings indicate that, while current models cannot meet the media industry's accuracy needs for full autonomy, they can serve as highly effective tools for enhancing human productivity. We conclude that a human-in-the-loop (HITL) approach is crucial and present the production-grade, cloud-based infrastructure we designed to support this workflow.

语音识别字幕生成人机协同意大利语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。