测试医学文本水印方法对事实性的影响,发现现有方法会破坏关键医疗信息
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
- 构建医学领域专用评估流程,兼顾事实准确性和连贯性
- 水印导致低熵医学术语被重写,事实错误率显著上升
- 适合关注医疗AI安全与内容可信度的研究者和开发者
随着大语言模型在医学等敏感领域的应用,其生成文本的流畅性带来了溯源与责任风险。水印技术通过嵌入可检测模式来缓解此类风险,但其在医学场景下的可靠性尚未验证。现有基准主要关注检测效果与流畅性的权衡,忽略了事实性风险。在医学文本中,水印常重分配低熵词元(高度可预测且多为关键医疗术语),这种调整可能导致事实偏差与幻觉,而通用基准未能捕捉此类问题。本文提出面向医学领域的评估流程,结合GPT-Judger与人工验证,引入事实加权得分(FWS)作为核心指标,强调事实准确性优先于连贯性。评估结果表明,当前水印方法显著损害医学事实性,熵值调整削弱了医学实体表达。研究呼吁开发注重领域特性的水印方案,以保障医学内容完整性。
原文摘要 · Abstract (English)
As large language models (LLMs) are adapted to sensitive domains such as medicine, their fluency raises safety risks, particularly regarding provenance and accountability. Watermarking embeds detectable patterns to mitigate these risks, yet its reliability in medical contexts remains untested. Existing benchmarks focus on detection-quality tradeoffs and overlook factual risks. In medical text, watermarking often reweights low-entropy tokens, which are highly predictable and often carry critical medical terminology. Shifting these tokens can cause inaccuracy and hallucinations, risks that prior general-domain benchmarks fail to capture. We propose a medical-focused evaluation workflow that jointly assesses factual accuracy and coherence. Using GPT-Judger and further human validation, we introduce the Factuality-Weighted Score (FWS), a composite metric prioritizing factual accuracy beyond coherence to guide watermarking deployment in medical domains. Our evaluation shows current watermarking methods substantially compromise medical factuality, with entropy shifts degrading medical entity representation. These findings underscore the need for domain-aware watermarking approaches that preserve the integrity of medical content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。