arXiv:2501.14288cs.CLcs.AI2025-01被引 3

用深度模型分析人机文本语义差异,提升检测准确性。

A Comprehensive Framework for Semantic Similarity Analysis of Human and AI-Generated Text Using Transformer Architectures and Ensemble Techniques

  • 融合DeBERTa与双向LSTM捕捉文本的局部和全局语义
  • 通过上下文增强与输出配置优化,识别率显著提升
  • 适合需要区分人机生成文本的场景,如内容审核

大语言模型的快速发展使检测AI生成文本成为紧迫挑战。传统方法难以捕捉人类与机器生成内容之间的细微语义差异。为此,我们提出一种基于语义相似性分析的新方法,采用多层架构结合预训练的DeBERTa-v3-large模型、双向LSTM和线性注意力池化,以捕捉局部与全局语义模式。为提升性能,引入领域级上下文融合与宽输出配置等先进输入输出增强技术,使模型能学习更具判别性的特征,并在不同领域间实现良好泛化。实验表明,该方法优于传统手段,证明其在AI生成文本检测及其他文本对比任务中的有效性。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has made detecting AI-generated text an increasingly critical challenge. Traditional methods often fail to capture the nuanced semantic differences between human and machine-generated content. We therefore propose a novel approach based on semantic similarity analysis, leveraging a multi-layered architecture that combines a pre-trained DeBERTa-v3-large model, Bi-directional LSTMs, and linear attention pooling to capture both local and global semantic patterns. To enhance performance, we employ advanced input and output augmentation techniques such as sector-level context integration and wide output configurations. These techniques enable the model to learn more discriminative features and generalize across diverse domains. Experimental results show that this approach works better than traditional methods, proving its usefulness for AI-generated text detection and other text comparison tasks.

文本检测语义分析DeBERTa

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。