arXiv:2608.20047cs.CLcs.CR2026-08

跨语言评估水印公平性,发现语言结构差异是核心问题

Auditing Cross-Lingual Fairness in Language Model Watermarking

论文配图:Auditing Cross-Lingual Fairness in Language Model Watermarking
图 1 · 摘自论文原文
  • 按部署场景校准检测阈值,避免单语言评估偏差
  • 跨语言水印检测失败率在不同语言族间差异显著
  • 适合关注多语言模型公平性与安全性的研究者

大型语言模型输出的水印方案几乎仅在英文上评估,依赖单一检测阈值和有限的质量指标。多语言部署暴露了这些设计选择在英文中无关紧要,但在跨语言时却决定结论的问题。本文提出一个包含四个部分的评估框架:按部署上下文经验校准检测阈值、不依赖阈值的辅助度量以区分校准失败与检测失败、三种独立的质量测量范式(分布、配对语义、参考困惑度),以及基于类型学家族划分的广义熵分解以分析跨语言差异。该框架应用于六种水印方案、三种开源生成器、覆盖四种书写系统和八个类型学家族的十一种语言,涵盖基础和指令微调两种模式,揭示了单一语言单一范式评估无法发现的失效模式。在检测与质量方面,观察到的差异主要存在于类型学家族之间,表明水印公平性差距本质上源于语言属性的结构性差异,而非特定语言的偶然特性。

原文摘要 · Abstract (English)

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

水印评估跨语言公平性语言结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。