构建首个学生同传数据集,支持细粒度对齐分析。
MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
- 收集7小时多语种同传录音,实现词与片段级对齐。
- 提出新评估指标与自动对齐基线方法。
- 开源工具InterAlign,支持长文本动态标注。
同传中译者需在源语句未完成时即开始翻译,延迟极短。为自动化理解与复现这一动态复杂任务,亟需专用数据集与分析工具,如平行语料库及自动标注工具。现有平行语料与对齐算法难以捕捉语段间的长程依赖关系,也无法建模原文与译文间的特定差异(如压缩、简化、功能泛化)。本文介绍MockConf:一个从学生课程模拟会议中采集的同传数据集,包含5种欧洲语言共7小时录音,已实现词级与片段级转录与对齐。我们还开发并发布了InterAlign——一款适用于长输入的现代网页版标注工具,支持平行词与片段标注,专为同传对齐设计。此外,提出了评估指标与自动对齐基线方法。数据集与工具已向社区公开。
原文摘要 · Abstract (English)
In simultaneous interpreting, an interpreter renders a source speech into another language with a very short lag, much sooner than sentences are finished. In order to understand and later reproduce this dynamic and complex task automatically, we need dedicated datasets and tools for analysis, monitoring, and evaluation, such as parallel speech corpora, and tools for their automatic annotation. Existing parallel corpora of translated texts and associated alignment algorithms hardly fill this gap, as they fail to model long-range interactions between speech segments or specific types of divergences (e.g., shortening, simplification, functional generalization) between the original and interpreted speeches. In this work, we introduce MockConf, a student interpreting dataset that was collected from Mock Conferences run as part of the students' curriculum. This dataset contains 7 hours of recordings in 5 European languages, transcribed and aligned at the level of spans and words. We further implement and release InterAlign, a modern web-based annotation tool for parallel word and span annotations on long inputs, suitable for aligning simultaneous interpreting. We propose metrics for the evaluation and a baseline for automatic alignment. Dataset and tools are released to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。