检测工具反而诱导用户多用LLM,且降低输出质量。
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

- 构建用户策略模型,分析检测如何改变使用与修改行为
- 检测导致用户使用率上升,输出质量反而下降
- 实证验证关键词频率呈现先升后降的反直觉趋势
随着大语言模型(LLM)应用普及,基于语言模式的检测工具和识别方法日益受到关注。这些检测器作为干预手段,不仅影响被检测属性,还作用于下游指标如模型使用率与输出质量。本文揭示了不完美的检测器会扭曲用户的使用激励机制,导致反直觉结果。我们提出一个简化模型,刻画用户如何战略性地决定使用量及内容后处理方式以规避检测。结果显示,检测可能反而促使用户增加使用频率;即使降低检测特征能提升质量,引入检测仍可能导致输出整体质量下降。我们还实证复现了在arXiv摘要中词频变化的‘先升后降’模式,说明检测干预可能引发复杂非线性反馈。本研究揭示了检测系统作为干预时对下游行为的潜在破坏性,暴露了其运行中的失效机制。
原文摘要 · Abstract (English)
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。