arXiv:2508.20718cs.CL2025-08EMNLP被引 3

解决大模型文本隐写与水印中的分词不一致问题

Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models

  • 发现导致分词不一致的异常词具频率低、临时性特征
  • 提出渐进验证和事后回滚方法,提升隐写流畅性与水印鲁棒性
  • 适合关注生成内容安全与对抗攻击的研究者

大语言模型显著提升了文本生成的质量与效率。一方面增强了文本隐写能力,另一方面也凸显了水印作为防范恶意滥用的重要手段。本文聚焦隐写与水印中通信双方Alice与Bob间的分词不一致(TI)问题,该问题会削弱系统鲁棒性。研究发现,造成TI的异常词具有两个关键特征:低频性和临时性。基于此,我们提出了两种针对性解决方案:面向隐写的渐进式验证方法,以及面向水印的事后回滚方法。实验表明:(1) 相较于传统消歧方法,直接解决TI可提升隐写的流畅性、不可察觉性及抗隐写分析能力;(2) 在水印任务中,解决TI能增强检测率并提高对攻击的鲁棒性。

原文摘要 · Abstract (English)

Large language models have significantly enhanced the capacities and efficiency of text generation. On the one hand, they have improved the quality of text-based steganography. On the other hand, they have also underscored the importance of watermarking as a safeguard against malicious misuse. In this study, we focus on tokenization inconsistency (TI) between Alice and Bob in steganography and watermarking, where TI can undermine robustness. Our investigation reveals that the problematic tokens responsible for TI exhibit two key characteristics: infrequency and temporariness. Based on these findings, we propose two tailored solutions for TI elimination: a stepwise verification method for steganography and a post-hoc rollback method for watermarking. Experiments show that (1) compared to traditional disambiguation methods in steganography, directly addressing TI leads to improvements in fluency, imperceptibility, and anti-steganalysis capacity; (2) for watermarking, addressing TI enhances detectability and robustness against attacks.

隐写水印大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。