用统一框架检测AI生成内容、幻觉等,效果优于传统方法。
A Unified Detection Framework for AI-Related Content and Artifacts

- 基于马氏距离构建统一检测框架,准确刻画正样本分布特征。
- 在多类正样本下实现高鲁棒性协方差估计,突破传统方法局限。
- 适用于文本生成、幻觉识别、水印检测等场景,适合安全监管应用。
人工智能既是利器也带来风险,有效监管需依赖对AI内容和伪造物的检测。本文提出一种基于马氏距离得分(MDS)的统一检测框架,适用于大语言模型生成文本、幻觉、水印及对抗样本等多种场景。核心在于准确建模正样本(如人类生成文本、真实陈述、无水印文本或非对抗样本)的分布,这需要高效稳健的深度表示协方差矩阵估计。针对正样本包含多类别且兼具同质与异质特性的问题,本文提出联合估计方案,同时优化逐样本和单元格级最小协方差确定(MCD)估计器,并设计高效优化算法,证明其收敛性。定义了联合估计器的破绽点并证明其具备高破绽点性质。实证评估验证了该框架的有效性。
原文摘要 · Abstract (English)
Artificial intelligence (AI) is a double-edged sword: while it has achieved remarkable success across a wide range of domains, its deployment also calls for effective oversight and regulation, for which the detection of AI-related content and artifacts is perhaps the most direct and cost-effective approach. To this end, we propose a unified detection framework based on Mahalanobis distance scores (MDS), applicable to several important settings, including the detection of large language model (LLM) generated text, hallucination, watermark, and adversarial examples. A key component of the proposed method is to accurately characterize the positive class--such as human-generated text, factual statements, unwatermarked text, or non-adversarial samples--which requires an efficient and robust estimator of the covariance matrix of deep representations of positive samples before computing the MDS. Since the positive samples typically consist of multiple classes, and these classes may exhibit both homogeneity and heterogeneity, we develop joint estimation methods for both the casewise and cellwise minimum covariance determinant (MCD) estimators. We provide efficient optimization algorithms for both estimators and prove their convergence. We provide a reasonable definition of the breakdown point for the joint estimators and prove their corresponding high breakdown point properties. Empirical evaluations confirm the effectiveness of the proposed detection framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。