通过分析文本风格稳定性,精准识别AI生成内容。
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation

- 结合离散风格特征与连续语义信号构建风格一致性画像。
- 跨领域检测性能领先现有方法11.79%,抗攻击能力更强。
- 适合需要高鲁棒性内容审核的平台和安全团队使用。
大型语言模型在内容生成中的广泛应用,使得区分人工撰写与模型生成文本成为多媒体内容审核的关键挑战。现有检测方法多依赖统计特征或模型特有启发式规则,易受改写和对抗攻击影响,导致鲁棒性和可解释性不足。本文提出LiSCP——一种轻量级风格一致性分析方法,聚焦于对抗性改写下的特征稳定性。该方法融合离散风格特征与连续语义信号,构建在多模态引导改写文本中保持稳定的风格一致性画像。在真实世界多媒体新闻、电影数据集及传统文本领域的实验表明,LiSCP在本域检测中表现优异,在跨域设置下性能超越现有方法最高达11.79%。同时在对抗攻击与人机混合场景下仍展现显著鲁棒性。
原文摘要 · Abstract (English)
The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from LLM-generated counterparts a critical task for multimedia moderation. Existing detectors often rely on statistical cues or model-specific heuristics, making them vulnerable to paraphrasing and adversarial manipulations, and consequently limiting their robustness and interpretability. In this work, we proposeLiSCP , a novel lightweight stylistic consistency profiling method for robust detection of LLM-generated textual content, focusing on feature stability under adversarial manipulation. Our approach constructs a consistency profile that combines discrete stylistic features with continuous semantic signals, leveraging stylistic stability across multimodal-guided paraphrased text variants. Experiments spanning real-world multimedia news and movie datasets and conventional text domains demonstrate that LiSCP achieves superior performance on in-domain detection and outperforms existing approaches by up to 11.79% in cross-domain settings. Additionally,it demonstrates notable robustness under adversarial scenarios, including adversarial attacks and hybrid human-AI settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。