新基准SHIELD评估大模型文本检测器的真实可靠性与稳定性。
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
- 引入可调难度的人类化框架,生成难检测的伪人类文本
- 提出融合可靠性与稳定性的统一评估指标,超越传统AUROC
- 适合关注实际部署中误报率和跨域鲁棒性的研究者
我们提出一种新型评估范式,用于衡量AI文本检测器在真实场景中的表现。现有方法多依赖于AUROC等常规指标,忽视了即使轻微的误报率也会严重阻碍检测系统的实际应用。此外,真实部署需要预设阈值,因此检测器在不同领域和对抗场景下保持性能稳定至关重要,而这一点在以往研究中被忽略。我们的基准SHIELD通过整合可靠性与稳定性因素,构建适用于实际评估的统一指标。同时,我们开发了一种后处理、模型无关的人类化框架,能可控地修改AI生成文本以更贴近人类写作风格。该难度感知方法有效挑战了当前最先进的零样本检测方法在可靠性与稳定性上的表现。(数据与代码:https://github.com/navid-aub/SHIELD-Benchmark)
原文摘要 · Abstract (English)
We present a novel evaluation paradigm for AI text detectors that prioritizes real-world and equitable assessment. Current approaches predominantly report conventional metrics like AUROC, overlooking that even modest false positive rates constitute a critical impediment to practical deployment of detection systems. Furthermore, real-world deployment necessitates predetermined threshold configuration, making detector stability (i.e. the maintenance of consistent performance across diverse domains and adversarial scenarios), a critical factor. These aspects have been largely ignored in previous research and benchmarks. Our benchmark, SHIELD, addresses these limitations by integrating both reliability and stability factors into a unified evaluation metric designed for practical assessment. Furthermore, we develop a post-hoc, model-agnostic humanification framework that modifies AI text to more closely resemble human authorship, incorporating a controllable hardness parameter. This hardness-aware approach effectively challenges current SOTA zero-shot detection methods in maintaining both reliability and stability. (Data and code: https://github.com/navid-aub/SHIELD-Benchmark)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。