arXiv:2506.09975cs.CL2025-06ACL被引 4

微调模型生成的社交媒体文本难被检测,威胁真实传播。

When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text

  • 用多种大模型生成50万条带争议话题的社交文本
  • 真实攻击者不公开微调模型时,检测率骤降至不足30%
  • 适合关注虚假信息防御与模型安全的研究者

检测社交媒体上的AI生成文本本已困难,因内容短、语言随意,更因该平台是网络影响力攻击的主要渠道。我们以中等实力攻击者视角,构建包含11个争议话题的505,159条由开源、闭源及微调大模型生成的文本数据集。在可访问生成模型的前提下,文本可被有效检测;但若攻击者仅使用未公开的微调模型,检测性能大幅下降。人类评估实验验证了该结果。消融实验证明多种检测算法对微调模型高度脆弱。该发现适用于所有检测场景,因微调是大模型常见且现实的应用方式。

原文摘要 · Abstract (English)

Detecting AI-generated text is a difficult problem to begin with; detecting AI-generated text on social media is made even more difficult due to the short text length and informal, idiosyncratic language of the internet. It is nonetheless important to tackle this problem, as social media represents a significant attack vector in online influence campaigns, which may be bolstered through the use of mass-produced AI-generated posts supporting (or opposing) particular policies, decisions, or events. We approach this problem with the mindset and resources of a reasonably sophisticated threat actor, and create a dataset of 505,159 AI-generated social media posts from a combination of open-source, closed-source, and fine-tuned LLMs, covering 11 different controversial topics. We show that while the posts can be detected under typical research assumptions about knowledge of and access to the generating models, under the more realistic assumption that an attacker will not release their fine-tuned model to the public, detectability drops dramatically. This result is confirmed with a human study. Ablation experiments highlight the vulnerability of various detection algorithms to fine-tuned LLMs. This result has implications across all detection domains, since fine-tuning is a generally applicable and realistic LLM use case.

AI检测大模型安全社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。