用思维链内部结构嵌入水印,防篡改且不破坏推理能力
Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought

- 通过对齐关键结构锚点与私有信号子空间实现水印嵌入
- 在微调、量化等攻击下仍能稳定检测,准确率超90%
- 适合保护具有复杂推理能力的大模型知识产权
具备思维链推理能力的大语言模型是重要知识产权,但现有黑盒水印方法常因扰动最终答案或依赖脆弱触发模式而牺牲推理准确性。本文提出BiCoT框架,通过将高显著性结构锚点与私有签名子空间对齐,并正则化普通控制令牌,将水印嵌入推理轨迹的内部几何结构中,使水印与推理相关表征耦合,移除需破坏推理特征。为应对模型被盗和表征漂移,引入基于Top-logprob的黑盒验证器RSR,利用哨兵令牌校准输出分布系统性偏移。实验表明,BiCoT在多种复杂推理任务中保持推理保真度,且在微调、量化、模型级扰动及自适应输出级攻击下均实现稳健检测,涵盖域内与域外场景。
原文摘要 · Abstract (English)
Large Language Models with Chain-of-Thought reasoning capabilities represent valuable intellectual property, yet existing black-box watermarking methods often trade robustness for reasoning fidelity by perturbing final answers or relying on fragile trigger patterns. We propose BiCoT, a watermarking framework that embeds ownership signals into the internal geometry of reasoning traces by aligning high-saliency structural anchors with a private signature subspace while regularizing ordinary control tokens to preserve semantic capacity. This design couples the watermark with reasoning-relevant representations, making removal difficult without disrupting the features that support coherent reasoning. To enable verification under model theft and representation drift, we introduce Robust Subspace Registration (RSR), a Top- logprob-based black-box verifier that uses sentinel tokens to calibrate systematic shifts in the output distribution. Experiments show that BiCoT preserves reasoning fidelity across diverse complex reasoning tasks while achieving robust detection under fine-tuning, quantization, model-level perturbations, and adaptive output-level attacks across in-domain and out-of-distribution settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。