给游戏AI加隐藏水印,既能防作弊又几乎不影响实力
Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games
- 将大模型水印技术移植到棋类游戏策略中
- 水印使策略性能下降极小,实验中可忽略不计
- 仅需几局对弈就能检测出水印,适合反作弊场景
大型语言模型的水印技术通过在输出中嵌入隐藏信息以验证来源,近年来备受关注,因其能识别模型的意外或故意滥用。类似问题也存在于游戏领域,例如检测在线象棋平台中未经授权使用AI工具的行为(如作弊)。本文首次研究游戏策略的水印方法,展示了如何将适用于大模型的KGW水印技术适配至完美信息扩展型博弈中的游戏策略。水印可通过统计检验进行检测。我们证明,水印对策略质量的影响(以期望效用衡量)可被控制,但存在可检测性与策略质量之间的权衡。实验中,我们将该框架应用于多个国际象棋引擎,结果表明:a) 水印对策略质量的影响可忽略;b) 仅需少量对局即可成功检测水印。
原文摘要 · Abstract (English)
Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained significant attention in recent days, thanks to their potential capability to detect accidental or deliberate misuse. Similar challenges involving model misuse also exist in the context of game-playing, such as when detecting the unauthorized use of AI tools in gaming platforms (e.g., cheating in online chess). In this paper, we initiate the study of how game-playing strategies can be watermarked. We show how the KGW watermark for LLMs can be adapted to watermark game-playing agents in perfect-information extensive-form games. The watermark can then be detected using a statistical test. We show that the degradation in the quality of the watermarked strategy profile, quantified by the expected utility, can be bounded, but there is a tradeoff between detectability and quality. In our experiments, we bootstrap the watermarking framework to various chess engines and demonstrate that a) the impact of the watermark on the quality of the strategy is negligible and b) the watermark can be detected with just a handful of games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。