用多智能体协作对抗机制,精准识别大模型生成文本
CAMF: Collaborative Adversarial Multi-agent Framework for Machine Generated Text Detection
- 多个大模型智能体分三阶段协同分析文本风格、语义与逻辑
- 在零样本场景下检测准确率显著超越现有方法
- 适合需要防伪和学术诚信保障的研究者使用
当前大型语言模型生成的机器文本检测面临日益严峻的挑战,涉及虚假信息传播与学术诚信风险。现有零样本检测方法普遍存在局限:(1)仅关注有限文本特征,分析浅显;(2)缺乏对风格、语义与逻辑等多维度一致性问题的深入探究。为此,我们提出协同对抗多智能体框架(CAMF),采用多个基于大模型的智能体,通过三阶段协同流程——多维语言特征提取、对抗一致性探查、综合判断聚合——实现对非人类生成文本中细微跨维度不一致性的深度分析。实证评估显示,CAMF在零样本条件下显著优于当前最先进的机器生成文本检测技术。
原文摘要 · Abstract (English)
Detecting machine-generated text (MGT) from contemporary Large Language Models (LLMs) is increasingly crucial amid risks like disinformation and threats to academic integrity. Existing zero-shot detection paradigms, despite their practicality, often exhibit significant deficiencies. Key challenges include: (1) superficial analyses focused on limited textual attributes, and (2) a lack of investigation into consistency across linguistic dimensions such as style, semantics, and logic. To address these challenges, we introduce the \textbf{C}ollaborative \textbf{A}dversarial \textbf{M}ulti-agent \textbf{F}ramework (\textbf{CAMF}), a novel architecture using multiple LLM-based agents. CAMF employs specialized agents in a synergistic three-phase process: \emph{Multi-dimensional Linguistic Feature Extraction}, \emph{Adversarial Consistency Probing}, and \emph{Synthesized Judgment Aggregation}. This structured collaborative-adversarial process enables a deep analysis of subtle, cross-dimensional textual incongruities indicative of non-human origin. Empirical evaluations demonstrate CAMF's significant superiority over state-of-the-art zero-shot MGT detection techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。