arXiv:2607.20462cs.AI2026-07

测试发现医学文本水印会引发严重错误,影响临床可靠性。

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

论文配图:Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
图 1 · 摘自论文原文
  • 首次系统评估5种水印对11个大模型在医学任务中的影响
  • 水印导致术语错乱、幻觉生成和图像描述遗漏等严重问题
  • 强调必须用专业医生验证的评估方式,避免临床风险被掩盖

大型语言模型(LLMs)正越来越多地融入临床工作流程,因此需要可靠的模型输出溯源机制,水印技术应运而生。然而,现有水印大多在通用基准上评估,未充分考察医学领域——该领域中微小的词汇级扰动可能引发显著语义变化。本文首次系统研究了水印对医学性能的影响,对5种水印方案在11个LLMs和7个VLMs上的多种单模态与多模态临床推理任务进行了评估。我们引入了一套由人类专家验证的流水线,用于系统审计医学推理质量、术语精确性及诱发的幻觉。结果表明,水印可能导致多项失效模式的显著退化,包括词汇扭曲、虚构术语和图像发现的误归因或遗漏。尤为重要的是,缺乏领域特异性分析,以及依赖掩盖临床文本固有缺陷的聚合指标,会系统性掩盖水印带来的实际性能下降。研究强调,在医学领域安全部署水印模型前,必须进行领域特定评估,否则现有基准可能隐藏具有临床后果的失败。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, underexplored. In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning. Importantly, we complement existing evaluations by introducing a human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations. Our results reveal that watermarking can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings. Notably, we find that the absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to clinical text, can systematically obscure practical watermark-induced degradations. Our findings establish domain-specific evaluation as a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures.

LLM水印医学AI模型可靠性临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。