arXiv:2601.07291cs.CVcs.AI2026-01被引 3

为大模型生成视觉语义自适应水印,保真度更高且不拖慢推理

A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model

  • 用轻量前缀调优提取视觉证据权重,动态选择可嵌入水印的词
  • 在视觉支持的词上集中加水印,提升视觉一致性7.8%(Chair-I)
  • 检测准确率96.88%,抗攻击能力99.3%,适合高保真多模态应用

水印已成为大视觉语言模型(LVLMs)内容溯源与知识产权保护的关键方案。然而,传统视觉无关水印会引入无关词汇,破坏视觉对齐;部分语义感知方法因拒绝采样导致推理延迟过高。本文提出视觉语义自适应水印(VISA-Mark),通过轻量级前缀调优动态提取视觉证据权重,量化候选词在视觉输入中的支持程度,并据此实现自适应词汇划分与逻辑值扰动,仅在视觉支持的词上强化水印。该方法显著提升视觉保真度,实验显示在Chair-I数据集上视觉一致性提升7.8%,同时保持96.88% AUC检测准确率和99.3%抗攻击鲁棒性,兼具高效性与可靠性,确立了多模态水印新标准。

原文摘要 · Abstract (English)

Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in Large Vision-Language Models (LVLMs). However, vision-agnostic watermarks introduce visually irrelevant tokens and disrupt visual grounding by enforcing indiscriminate pseudo-random biases, while some semantic-aware methods incur prohibitive inference latency due to rejection sampling. In this paper, we propose the VIsual Semantic Adaptive Watermark (VISA-Mark), a novel framework that embeds detectable signals while strictly preserving visual fidelity. Our approach employs a lightweight, efficiently trained prefix-tuner to extract dynamic Visual-Evidence Weights, which quantify the evidentiary support for candidate tokens based on the visual input. These weights guide an adaptive vocabulary partitioning and logits perturbation mechanism, concentrating watermark strength specifically on visually-supported tokens. By actively aligning the watermark with visual evidence, VISA-Mark effectively maintains visual fidelity. Empirical results confirm that VISA-Mark outperforms conventional methods with a 7.8% improvement in visual consistency (Chair-I) and superior semantic fidelity. The framework maintains highly competitive detection accuracy (96.88% AUC) and robust attack resilience (99.3%) without sacrificing inference efficiency, effectively establishing a new standard for reliability-preserving multimodal watermarking.

水印视觉语言模型前缀调优保真度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。