构建首个多语言实用显化语料库,助力机器翻译理解文化隐含信息。
PETra: A Multilingual Corpus of Pragmatic Explicitation in Translation
- 基于空对齐与主动学习,自动识别跨语言显化现象。
- 实体和系统级显化最常见,主动学习提升准确率7-8个百分点。
- 适合研究机器翻译中的文化适配与跨语言语用学。
译者常通过补充背景细节,使原文隐含的文化含义在目标语中明确。这种现象称为语用显化,在翻译理论中广受讨论,但极少被计算建模。本文提出PragExTra,首个多语言语用显化语料库与检测框架,涵盖来自TED-Multi和Europarl的八组语言对,包含实体描述、单位换算及译者注释等新增内容。通过空对齐识别候选案例,并结合人工标注进行主动学习优化。结果表明,实体与系统级显化最为频繁,主动学习使分类器准确率提升7-8个百分点,跨语言最高达0.88准确率与0.82 F1值。PragExTra证实语用显化是可度量的跨语言现象,为构建具备文化感知能力的机器翻译迈出关键一步。
原文摘要 · Abstract (English)
Translators often enrich texts with background details that make implicit cultural meanings explicit for new audiences. This phenomenon, known as pragmatic explicitation, has been widely discussed in translation theory but rarely modeled computationally. We introduce PragExTra, the first multilingual corpus and detection framework for pragmatic explicitation. The corpus covers eight language pairs from TED-Multi and Europarl and includes additions such as entity descriptions, measurement conversions, and translator remarks. We identify candidate explicitation cases through null alignments and refined using active learning with human annotation. Our results show that entity and system-level explicitations are most frequent, and that active learning improves classifier accuracy by 7-8 percentage points, achieving up to 0.88 accuracy and 0.82 F1 across languages. PragExTra establishes pragmatic explicitation as a measurable, cross-linguistic phenomenon and takes a step towards building culturally aware machine translation. Keywords: translation, multilingualism, explicitation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。