无需重新训练,即可让文本重写检测器获得精确的假阳性控制。
A Distribution-Free Framework for Rewrite-Based Human-text Detection via Knockoff Filtering
- 利用反例构造思想,将检测问题转化为多重假设检验。
- 在3个模型、19个领域、4个大模型上实现稳定假阳性率控制。
- 适合需要可靠误报率保障的AI文本检测场景。
我们提出一种分布无关的统计框架,可在不重新训练的情况下,将任意基于重写的检测器转化为具有有限样本错误发现率(FDR)保证的检测器。核心观察是,重写检测隐式构建了反例样本,使大语言模型生成文本的检测可被形式化为具有反例结构的多重假设检验问题。这一视角将检测统计量设计与误报控制分离,使现有重写检测器仅通过简单校准即可继承有限样本下的FDR保证。我们在三个检测模型、19个应用领域和四个大语言模型上验证了该方法在保持有效检测力的同时,实现了可靠的FDR控制。
原文摘要 · Abstract (English)
We propose a distribution-free statistical framework that converts arbitrary rewrite-based detectors into detectors with finite-sample FDR guarantees without retraining. Our key observation is that rewrite-based detection implicitly constructs knockoff samples, enabling LLM-generated text detection to be formulated as a multiple hypothesis testing problem with knockoff structure. This perspective separates the design of detection statistics from the control of false discoveries, allowing existing rewrite detectors to inherit finite-sample false discovery rate (FDR) guarantees through a simple calibration procedure. We demonstrate reliable FDR control with meaningful detection power across three detection models, 19 domains, and four LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。