多语言模型评估中,形式与语义差异导致现有指标不可靠。
Form and Meaning in Intrinsic Multilingual Evaluations
- 揭示多语言评估中隐含的假设:平行句语义相同
- 六种指标在双语语料上表现不一致,无法通用比较
- 提出形式与意义分离问题,警示评估设计者
条件语言模型的内在评估指标(如困惑度、每字符比特数)在单语和多语场景中广泛使用。单语设置下这些指标简单易用且可比,但在多语场景中依赖若干假设。其中关键假设是:在平行句上比较模型困惑度能反映其质量,因信息内容(即语义)相同。然而,这些指标本质上衡量的是信息论意义上的信息量,而非语义。本文明确指出此类假设,并分析其影响。我们在两个多语平行语料库上,对六种指标进行了实验,涵盖单语与多语模型。结果表明,当前指标不具备普遍可比性。通过探讨形式与意义之争,为该现象提供解释。
原文摘要 · Abstract (English)
Intrinsic evaluation metrics for conditional language models, such as perplexity or bits-per-character, are widely used in both mono- and multilingual settings. These metrics are rather straightforward to use and compare in monolingual setups, but rest on a number of assumptions in multilingual setups. One such assumption is that comparing the perplexity of CLMs on parallel sentences is indicative of their quality since the information content (here understood as the semantic meaning) is the same. However, the metrics are inherently measuring information content in the information-theoretic sense. We make this and other such assumptions explicit and discuss their implications. We perform experiments with six metrics on two multi-parallel corpora both with mono- and multilingual models. Ultimately, we find that current metrics are not universally comparable. We look at the form-meaning debate to provide some explanation for this.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。