arXiv:2507.20762cs.IR2025-07被引 1

给大模型生成的时序数据加数字水印,防抄袭又防造假。

Watermarking Large Language Model-based Time Series Forecasting

  • 利用时序片段与冷词嵌入的相似性差异,嵌入不可察觉的水印信号。
  • 在7个数据集上检测准确率超95%,生成数据质量下降不足1%。
  • 适用于各类时序大模型,适合关注版权与安全的开发者。

基于大语言模型的时序预测(LLMTS)在处理复杂多样的时间序列数据方面展现出显著潜力,是迈向时序分析基础模型的重要一步。然而,这一新兴范式带来两个关键挑战:其一,巨大的商业价值和资源密集型开发引发知识产权保护的紧迫需求;其二,其强大的预测能力可能被滥用,生成误导性或伪造的深度伪造时序数据。为应对这些问题,我们探索对LLMTS模型输出进行水印标记,即在生成的时序数据中嵌入难以察觉但可由专用算法检测的信号。我们提出一种新型后置水印框架Waltz,广泛兼容现有LLMTS模型。Waltz受启发于一个经验观察:时序片段嵌入极少与特定的一组大模型词元对齐,我们称之为“冷词”。利用这一特性,Waltz通过重连片段嵌入与冷词嵌入之间的相似性统计来嵌入水印,并使用相似性z分数进行检测。为最小化潜在副作用,我们引入基于相似性的嵌入位置识别策略,并采用投影梯度下降将水印噪声约束在指定范围内。在两个主流LLMTS模型上,针对七个基准数据集的大量实验表明,Waltz在保持生成数据质量几乎不变的前提下,实现了超过95%的水印检测准确率。

原文摘要 · Abstract (English)

Large Language Model-based Time Series Forecasting (LLMTS) has shown remarkable promise in handling complex and diverse temporal data, representing a significant step toward foundation models for time series analysis. However, this emerging paradigm introduces two critical challenges. First, the substantial commercial potential and resource-intensive development raise urgent concerns about intellectual property (IP) protection. Second, their powerful time series forecasting capabilities may be misused to produce misleading or fabricated deepfake time series data. To address these concerns, we explore watermarking the outputs of LLMTS models, that is, embedding imperceptible signals into the generated time series data that remain detectable by specialized algorithms. We propose a novel post-hoc watermarking framework, Waltz, which is broadly compatible with existing LLMTS models. Waltz is inspired by the empirical observation that time series patch embeddings are rarely aligned with a specific set of LLM tokens, which we term ``cold tokens''. Leveraging this insight, Waltz embeds watermarks by rewiring the similarity statistics between patch embeddings and cold token embeddings, and detects watermarks using similarity z-scores. To minimize potential side effects, we introduce a similarity-based embedding position identification strategy and employ projected gradient descent to constrain the watermark noise within a defined boundary. Extensive experiments using two popular LLMTS models across seven benchmark datasets demonstrate that Waltz achieves high watermark detection accuracy with minimal impact on the quality of the generated time series.

时序预测数字水印大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。