arXiv:2505.05084cs.CL2025-05ACL被引 11

用多尺度方法在零样本下精准控制误报率,提升文本生成检测可靠性。

Reliably Bounding False Positives: A Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction

  • 引入多尺度置信预测框架,在不依赖训练数据的前提下约束误报率。
  • 在多个数据集上实现高检测率且误报率低于设定阈值,对抗攻击下仍稳定。
  • 配合高质量真实文本数据集RealDet,适合需高可信度检测的场景。

大语言模型的快速发展引发了恶意使用风险,开发有效检测工具已成为当务之急。然而,现有方法过度关注检测准确率,常忽视高误报率带来的社会风险。本文利用置信预测(Conformal Prediction, CP)技术,有效控制误报率上限。直接应用CP虽能约束误报率,但会显著降低检测性能。为此,本文提出基于多尺度置信预测(Multiscaled Conformal Prediction, MCP)的零样本机器生成文本检测框架,既满足误报率约束,又提升检测性能。同时引入RealDet数据集,覆盖广泛领域,确保校准真实性和检测效果。实证结果表明,MCP能有效控制误报率,显著提升检测性能,并增强对对抗攻击的鲁棒性,适用于多种检测器与数据集。

原文摘要 · Abstract (English)

The rapid advancement of large language models has raised significant concerns regarding their potential misuse by malicious actors. As a result, developing effective detectors to mitigate these risks has become a critical priority. However, most existing detection methods focus excessively on detection accuracy, often neglecting the societal risks posed by high false positive rates (FPRs). This paper addresses this issue by leveraging Conformal Prediction (CP), which effectively constrains the upper bound of FPRs. While directly applying CP constrains FPRs, it also leads to a significant reduction in detection performance. To overcome this trade-off, this paper proposes a Zero-Shot Machine-Generated Text Detection Framework via Multiscaled Conformal Prediction (MCP), which both enforces the FPR constraint and improves detection performance. This paper also introduces RealDet, a high-quality dataset that spans a wide range of domains, ensuring realistic calibration and enabling superior detection performance when combined with MCP. Empirical evaluations demonstrate that MCP effectively constrains FPRs, significantly enhances detection performance, and increases robustness against adversarial attacks across multiple detectors and datasets.

文本检测置信预测零样本误报控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。