arXiv:2606.17354cs.CLcs.AI2026-06

构建可操作的不可译性框架,助力机器翻译应对跨语言意义丢失问题。

Translating the Untranslatable: An Operationalizable Ontology for Untranslatability

论文配图:Translating the Untranslatable: An Operationalizable Ontology for Untranslatability
图 1 · 摘自论文原文
  • 提出不可译性分类体系与补偿策略,明确如何传递跨语言意义。
  • 构建多语言不可译句子数据集,包含策略导向的翻译对。
  • 实验证明带解释性上下文的标注补偿策略更受人类青睐。

不可译性指语义在跨语言转换中无法直接保留的情况,在语言学中已有研究,但在自然语言处理领域仍属空白。随着机器翻译系统在标准基准上表现提升,其局限性日益集中在这些无法实现一一对应的情形。本文提出一种结构化的不可译性本体论及补偿策略分类体系,补偿策略是在不可译情境下传递意义的具体方法。我们将该框架操作化为一个多语言不可译句子数据集,每条句子配有基于策略的翻译,支持对翻译行为的受控分析。初步的人类偏好研究表明,翻译质量取决于所用策略,且一致偏好包含解释性上下文的“标注补偿”策略。该框架与数据集为研究和建模策略导向的机器翻译奠定了基础。

原文摘要 · Abstract (English)

Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in NLP. As machine translation (MT) systems improve on standard benchmarks, their limitations increasingly concentrate in such cases, where translation cannot be reduced to one-to-one equivalence. We introduce a structured ontology of untranslatability along with a taxonomy of compensation strategies, which are specific techniques to convey meaning under these untranslatable circumstances. We operationalize this framework into a multilingual dataset of untranslatable sentences paired with strategy-based translations, enabling controlled analysis of translation behavior. Initial human preference studies suggest that translation quality depends on the strategy used, with consistent preferences for outputs that include explanatory context, known as the Annotation compensation strategy. Our framework and dataset provide a foundation for studying and modeling strategy-informed machine translation.

机器翻译不可译性语义保留策略框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。