提出温度调控水印方法,提升对抗改写攻击的鲁棒性。
Temperature Matters: Enhancing Watermark Robustness Against Paraphrasing Attacks
- 通过调节生成温度增强水印信号稳定性。
- 在改写文本测试中,水印保留率显著高于基线方法。
- 适合关注AI内容溯源与伦理安全的研究者。
当前大型语言模型(LLMs)已广泛应用于社会各领域,其强大能力也带来滥用风险。为此,学术界提出在机器生成文本中嵌入水印标记以实现算法识别。本文首先复现了先前基线研究,发现其对生成模型变化敏感。随后提出一种新型水印方法,并通过改写生成文本进行严格评估。实验结果表明,该方法在对抗改写攻击时表现出更强的鲁棒性,显著优于~\cite{aarson} 的水印方案。
原文摘要 · Abstract (English)
In the present-day scenario, Large Language Models (LLMs) are establishing their presence as powerful instruments permeating various sectors of society. While their utility offers valuable support to individuals, there are multiple concerns over potential misuse. Consequently, some academic endeavors have sought to introduce watermarking techniques, characterized by the inclusion of markers within machine-generated text, to facilitate algorithmic identification. This research project is focused on the development of a novel methodology for the detection of synthetic text, with the overarching goal of ensuring the ethical application of LLMs in AI-driven text generation. The investigation commences with replicating findings from a previous baseline study, thereby underscoring its susceptibility to variations in the underlying generation model. Subsequently, we propose an innovative watermarking approach and subject it to rigorous evaluation, employing paraphrased generated text to asses its robustness. Experimental results highlight the robustness of our proposal compared to the~\cite{aarson} watermarking method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。