arXiv:2606.07219cs.CLcs.SI2026-06

用对抗方法生成并检测假社交账号内容,提升真实场景识别能力。

Adversarial Creation and Detection of AI-Generated Social Bot Content

论文配图:Adversarial Creation and Detection of AI-Generated Social Bot Content
图 1 · 摘自论文原文
  • 构建对抗性数据集,模拟恶意者伪装真人发帖
  • 在跨平台多语言数据上实现高精度检测
  • 适合安全研究与平台反作弊系统开发者

大型语言模型与社交机器人结合,使恶意行为者能大规模生成类人内容,扰乱信息生态。现有内容检测模型在真实场景中表现不佳,主要因缺乏真实标注数据。本文提出一种对抗性方法,模拟恶意者对真实社交媒体用户的模仿行为,构建了一个多语言、跨平台的成对人类与AI生成消息数据集。基于该对抗数据训练的模型,在真实世界、分布外的数据上显著优于现有内容型机器人检测模型。

原文摘要 · Abstract (English)

The convergence of large language models and social bots allows malicious actors to manipulate the information ecosystem by generating human-like content at scale. Existing models for detecting AI-generated content often fail in the wild, primarily due to the lack of ground-truth data. We address this gap through an adversarial methodology that models the impersonation of real social media users by malicious actors. Using this methodology, we curate a multilingual, cross-platform dataset of paired human and AI-generated messages. Training on such adversarial data yields accurate detection of AI-generated text. Our approach significantly outperforms existing models for content-based bot detection in real-world, out-of-distribution data.

社交机器人对抗生成内容检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。