arXiv:2503.15044cs.CL2025-03被引 1

用结构化提示生成对话数据,提升机器生成文本检测效果。

SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection

  • 设计提示框架,自动构造正负样本增强数据集
  • 构建14个新对话数据集,混合训练使检测准确率提升
  • 适合关注LLM安全与对话检测的研究者使用

大语言模型生成合成内容的能力日益增强,引发对其滥用的担忧,推动了机器生成文本(MGT)检测模型的发展。然而,这些检测器因缺乏高质量合成数据集而面临训练难题。为此,我们提出SPADE,一种基于提示的结构化框架,用于检测合成对话,通过生成正负样本实现数据增强。该方法生成14个新的对话数据集,并在8个MGT检测模型上进行基准测试。结果表明,使用所提出的混合数据集可显著提升模型泛化性能,为提升LLM应用安全性提供实用方案。考虑到真实场景中对话代理无法预知对方后续发言,我们模拟在线对话检测,分析了对话历史长度与检测准确率的关系。相关开源数据、代码与提示已发布于https://github.com/AngieYYF/SPADE-customer-service-dialogue。

原文摘要 · Abstract (English)

The increasing capability of large language models (LLMs) to generate synthetic content has heightened concerns about their misuse, driving the development of Machine-Generated Text (MGT) detection models. However, these detectors face significant challenges due to the lack of high-quality synthetic datasets for training. To address this issue, we propose SPADE, a structured framework for detecting synthetic dialogues using prompt-based positive and negative samples. Our proposed methods yield 14 new dialogue datasets, which we benchmark against eight MGT detection models. The results demonstrate improved generalization performance when utilizing a mixed dataset produced by proposed augmentation frameworks, offering a practical approach to enhancing LLM application security. Considering that real-world agents lack knowledge of future opponent utterances, we simulate online dialogue detection and examine the relationship between chat history length and detection accuracy. Our open-source datasets, code and prompts can be downloaded from https://github.com/AngieYYF/SPADE-customer-service-dialogue.

对话检测LLM安全数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。