arXiv:2603.09785cs.CL2026-03被引 1

整合英德双语语料,支持语言差异与口译研究

EPIC-EuroParl-UdS: Information-Theoretic Perspectives on Translation and Interpreting

  • 融合口语与书面语语料,新增词对齐与意外度指标
  • 验证重建口语数据可靠性,预测口译填充词准确率达72.3%
  • 适合研究语言变异、口译特点及翻译风格的学者

本文发布了一个更新整合的双向英德语语料库,包含欧洲议会原始演讲及其翻译与口译文本。新版本修正了先前使用中发现的元数据和文本错误,优化内容质量,更新语言标注,并新增词对齐与词级意外度指数等层级。该资源旨在支持信息论方法研究语言差异,尤其适用于对比书面与口语模式、分析言语不连贯现象,以及传统翻译风格研究(包括平行与可比分析)。论文概述本次更新内容,总结以往基于该语料的研究成果,并开展一项新示范研究:验证重建口语数据的完整性,并评估基于基础与微调GPT-2模型及机器翻译系统在口译填充词预测任务中的概率度量表现。

原文摘要 · Abstract (English)

This paper introduces an updated and combined version of the bidirectional English-German EPIC-UdS (spoken) and EuroParl-UdS (written) corpora containing original European Parliament speeches as well as their translations and interpretations. The new version corrects metadata and text errors identified through previous use, refines the content, updates linguistic annotations, and adds new layers, including word alignment and word-level surprisal indices. The combined resource is designed to support research using information-theoretic approaches to language variation, particularly studies comparing written and spoken modes, and examining disfluencies in speech, as well as traditional translationese studies, including parallel (source vs. target) and comparable (original vs. translated) analyses. The paper outlines the updates introduced in this release, summarises previous results based on the corpus, and presents a new illustrative study. The study validates the integrity of the rebuilt spoken data and evaluates probabilistic measures derived from base and fine-tuned GPT-2 and machine translation models on the task of filler particles prediction in interpreting.

语料库口译研究信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。