arXiv:2511.21009cs.LG2025-11

用XLM-RoBERTa检测AI生成文本,区分纯生成与改写内容。

ChatGpt Content detection: A new approach using xlm-roberta alignment

  • 基于XLM-RoBERTa构建多语言文本检测模型。
  • 在多种文体上实现高准确率,关键特征为困惑度与注意力机制。
  • 适合关注学术诚信与AI伦理的研究者使用。

随着ChatGPT等生成式AI技术日益普及,区分人工智能生成文本与人工撰写内容的挑战愈发紧迫。本文提出一种综合方法,用于检测完全由AI生成的文本以及经AI重写的真人文本。采用先进的多语言变压器模型XLM-RoBERTa,结合严格的预处理与特征提取,包括困惑度、语义和可读性特征。在平衡的人类与AI生成文本数据集上微调模型,并评估其性能。结果显示,该模型在各类文本体裁中均表现出高准确率和强鲁棒性。通过特征分析发现,困惑度与基于注意力的特征在区分人机文本中起关键作用。研究结果为维护学术诚信提供有力工具,推动人工智能伦理中的透明性与问责制。未来工作将探索更先进模型并扩展数据集以提升泛化能力。

原文摘要 · Abstract (English)

The challenge of separating AI-generated text from human-authored content is becoming more urgent as generative AI technologies like ChatGPT become more widely available. In this work, we address this issue by looking at both the detection of content that has been entirely generated by AI and the identification of human text that has been reworded by AI. In our work, a comprehensive methodology to detect AI- generated text using XLM-RoBERTa, a state-of-the-art multilingual transformer model. Our approach includes rigorous preprocessing, and feature extraction involving perplexity, semantic, and readability features. We fine-tuned the XLM-RoBERTa model on a balanced dataset of human and AI-generated texts and evaluated its performance. The model demonstrated high accuracy and robust performance across various text genres. Additionally, we conducted feature analysis to understand the model's decision-making process, revealing that perplexity and attention-based features are critical in differentiating between human and AI-generated texts. Our findings offer a valuable tool for maintaining academic integrity and contribute to the broader field of AI ethics by promoting transparency and accountability in AI systems. Future research directions include exploring other advanced models and expanding the dataset to enhance the model's generalizability.

文本检测XLM-RoBERTaAI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。