arXiv:2511.03641cs.CRcs.AI2025-11

为欧盟AI法案落地,提出LLM水印技术评估框架。

Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology

  • 按模型生命周期阶段分类水印技术,清晰划分应用节点。
  • 将法案四标准转化为可测指标,揭示现有方法均未达标。
  • 建议深入研究嵌入模型底层架构的水印方案。

为推动欧洲联盟内可信赖的人工智能发展,欧盟《人工智能法案》要求通用大模型提供者对其输出进行标记并可检测。第50条与第133条强调标记方法必须具备“足够可靠、互操作性、有效性和鲁棒性”。然而,大语言模型(LLMs)水印技术快速演进且类型多样,难以将上述四项标准转化为具体可衡量的评估标准。本文旨在将欧洲规范要求锚定于水印技术的多元实践之中。贡献有三:(1) 提出水印方法在模型生命周期中应用阶段的分类体系——训练前、训练中、训练后,以及采样或下一词分布阶段;(2) 通过映射当前最先进评估结果,解读法案中的四个标准,涵盖水印鲁棒性、可检测性及模型质量,并针对尚缺乏理论支撑的互操作性,提出三个规范性维度以指导评估;(3) 对比现有水印方法与已操作化的欧洲标准,发现目前无一方法满足全部四项要求。基于新兴实证测试,建议进一步研究直接嵌入大模型底层架构的水印技术。

原文摘要 · Abstract (English)

To foster trustworthy Artificial Intelligence (AI) within the European Union, the AI Act requires providers to mark and detect the outputs of their general-purpose models. The Article 50 and Recital 133 call for marking methods that are ''sufficiently reliable, interoperable, effective and robust''. Yet, the rapidly evolving and heterogeneous landscape of watermarks for Large Language Models (LLMs) makes it difficult to determine how these four standards can be translated into concrete and measurable evaluations. Our paper addresses this challenge, anchoring the normativity of European requirements in the multiplicity of watermarking techniques. Introducing clear and distinct concepts on LLM watermarking, our contribution is threefold. (1) Watermarking Categorisation: We propose an accessible taxonomy of watermarking methods according to the stage of the LLM lifecycle at which they are applied - before, during, or after training, and during next-token distribution or sampling. (2) Watermarking Evaluation: We interpret the EU AI Act's requirements by mapping each criterion with state-of-the-art evaluations on robustness and detectability of the watermark, and of quality of the LLM. Since interoperability remains largely untheorised in LLM watermarking research, we propose three normative dimensions to frame its assessment. (3) Watermarking Comparison: We compare current watermarking methods for LLMs against the operationalised European criteria and show that no approach yet satisfies all four standards. Encouraged by emerging empirical tests, we recommend further research into watermarking directly embedded within the low-level architecture of LLMs.

AI监管水印技术LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。