arXiv:2505.11550cs.CLcs.AI2025-05被引 11

提出两种模型,精准识别文本是否由AI生成并追溯来源模型。

AI-generated Text Detection: A Multifaceted Approach to Binary and Multiclass Classification

  • 设计优化与简化双架构,分别应对二分类与多分类任务。
  • 在二分类任务中达0.994 F1,多分类任务中达0.627 F1。
  • 适合关注AI文本检测与模型溯源的研究者与应用开发者。

大型语言模型(LLMs)在生成风格多样、接近人类写作的文本方面表现出色,但其能力可能被滥用于制造假新闻、垃圾邮件或学术不端。因此,准确检测AI生成文本并识别其来源模型,对确保LLMs的负责任使用至关重要。本文针对AAAI 2025年Defactify研讨会提出的文本生成检测共享任务中的两个子任务展开研究:任务A为区分人工撰写与AI生成文本,任务B为识别文本所属的语言模型。针对每项任务,我们提出了两种神经网络架构:一种优化模型与一种简化版本。在任务A中,优化模型取得第五名,F1得分为0.994;在任务B中,简化模型同样位列第五,F1得分为0.627。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in generating text that closely resembles human writing across a wide range of styles and genres. However, such capabilities are prone to potential misuse, such as fake news generation, spam email creation, and misuse in academic assignments. As a result, accurate detection of AI-generated text and identification of the model that generated it are crucial for maintaining the responsible use of LLMs. In this work, we addressed two sub-tasks put forward by the Defactify workshop under AI-Generated Text Detection shared task at the Association for the Advancement of Artificial Intelligence (AAAI 2025): Task A involved distinguishing between human-authored or AI-generated text, while Task B focused on attributing text to its originating language model. For each task, we proposed two neural architectures: an optimized model and a simpler variant. For Task A, the optimized neural architecture achieved fifth place with $F1$ score of 0.994, and for Task B, the simpler neural architecture also ranked fifth place with $F1$ score of 0.627.

文本检测模型溯源LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。