arXiv:2502.16857cs.CLcs.AI2025-02被引 4

通过加噪与DeBERTa集成模型,实现高精度检测AI生成文本。

Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models

  • 对数据添加噪声提升模型鲁棒性
  • 使用DeBERTa模型集成,F1达0.9531
  • 适合关注AI文本检测的从业者

本文提出一种针对AI生成文本的检测方法,用于第四届多模态事实核查与仇恨言论检测研讨会的Defactify 4.0共享任务。该任务包含两个子任务:Task-A为判断文本是否由AI生成;Task-B为识别具体生成该文本的大语言模型。我们团队(Sarang)在两项任务中均获第一名,对应F1分数分别为1.0和0.9531。方法上,通过向数据集添加噪声以增强模型的鲁棒性和泛化能力,并采用DeBERTa模型的集成策略,有效捕捉文本中的复杂模式。结果表明,基于噪声驱动与模型集成的方法在检测性能上表现优异,为未来AI生成内容检测提供了新范式。

原文摘要 · Abstract (English)

This paper presents an effective approach to detect AI-generated text, developed for the Defactify 4.0 shared task at the fourth workshop on multimodal fact checking and hate speech detection. The task consists of two subtasks: Task-A, classifying whether a text is AI generated or human written, and Task-B, classifying the specific large language model that generated the text. Our team (Sarang) achieved the 1st place in both tasks with F1 scores of 1.0 and 0.9531, respectively. The methodology involves adding noise to the dataset to improve model robustness and generalization. We used an ensemble of DeBERTa models to effectively capture complex patterns in the text. The result indicates the effectiveness of our noise-driven and ensemble-based approach, setting a new standard in AI-generated text detection and providing guidance for future developments.

文本检测DeBERTaAI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。