用语义水印标记学术审稿中的AI生成内容,确保可追溯又不降低质量。
The Feasibility of Topic-Based Watermarking on Academic Peer Reviews
- 在审稿文本中嵌入语义感知水印,保持原内容质量。
- 水印在改写后仍能稳定检测,准确率超90%。
- 适合需要监管AI使用但又不破坏审稿流程的期刊会议。
大型语言模型(LLMs)正越来越多地被用于学术工作流,如语言润色和文献摘要。然而,由于存在保密泄露、幻觉内容和评估不一致等风险,其在同行评审中的使用仍被禁止。随着LLM生成文本越来越接近人类写作,亟需可靠的溯源机制以维护评审过程的完整性。本文评估了一种基于主题的水印技术(Topic-Based Watermarking, TBW),该技术可将可检测信号嵌入生成文本中。我们在多个LLM配置下(包括基础模型、少样本和微调版本)使用真实学术会议的审稿数据进行了系统评估。结果表明,TBW在保持与非水印输出相当的评审质量的同时,在重述(paraphrasing)攻击下仍具备鲁棒的检测性能。这些发现证明了TBW作为最小侵入性且实用的解决方案,在同行评审中实现LLM溯源的可行性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly integrated into academic workflows, with many conferences and journals permitting their use for tasks such as language refinement and literature summarization. However, their use in peer review remains prohibited due to concerns around confidentiality breaches, hallucinated content, and inconsistent evaluations. As LLM-generated text becomes more indistinguishable from human writing, there is a growing need for reliable attribution mechanisms to preserve the integrity of the review process. In this work, we evaluate topic-based watermarking (TBW), a semantic-aware technique designed to embed detectable signals into LLM-generated text. We conduct a systematic assessment across multiple LLM configurations, including base, few-shot, and fine-tuned variants, using authentic peer review data from academic conferences. Our results show that TBW maintains review quality relative to non-watermarked outputs, while demonstrating robust detection performance under paraphrasing. These findings highlight the viability of TBW as a minimally intrusive and practical solution for LLM attribution in peer review settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。