arXiv:2603.16258cs.CL2026-03

用ASR辅助转录能提速,但准确率不稳,需结合经验与对话类型调整。

Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus

  • 对比人工与ASR辅助转录,测试不同对话类型和标注者经验的影响。
  • ASR可提升30%以上转录速度,但整体准确率提升不显著。
  • 适合需要快速构建语料库且有经验标注员的项目参考使用。

本文研究了自动语音识别(ASR)在意大利语口语语料库KIParla转录流程中的应用。通过两阶段实验,11名专家与新手标注员对三类对话音频分别进行人工与ASR辅助转录,并基于词级对齐、统计建模及注释评估指标进行分析。结果表明,ASR辅助可提升30%以上的转录速度,但整体准确率未稳定提高,效果受工作流配置、对话类型及标注者经验影响。结合对齐指标、描述性统计与统计建模的方法,为跨标注者与工作流的转录行为监控提供了系统框架。尽管存在局限,经任务特定微调后,ASR辅助转录或可有效加速语料库建设而不牺牲质量。

原文摘要 · Abstract (English)

This paper analyses the implementation of Automatic Speech Recognition (ASR) into the transcription workflow of the KIParla corpus, a resource of spoken Italian. Through a two-phase experiment, 11 expert and novice transcribers produced both manual and ASR-assisted transcriptions of identical audio segments across three different types of conversation, which were subsequently analyzed through a combination of statistical modeling, word-level alignment and a series of annotation-based metrics. Results show that ASR-assisted workflows can increase transcription speed but do not consistently improve overall accuracy, with effects depending on multiple factors such as workflow configuration, conversation type and annotator experience. Analyses combining alignment-based metrics, descriptive statistics and statistical modeling provide a systematic framework to monitor transcription behavior across annotators and workflows. Despite limitations, ASR-assisted transcription, potentially supported by task-specific fine-tuning, could be integrated into the KIParla transcription workflow to accelerate corpus creation without compromising transcription quality.

语音转录语料库构建ASR应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。