arXiv:2511.08455cs.CL2025-11AAAI

用大模型生成反事实数据,让社交机器人检测器更抗误导性文本陷阱。

Bot Meets Shortcut: How Can LLMs Aid in Handling Unknown Invariance OOD Scenarios?

  • 通过构造虚假文本关联测试检测器鲁棒性,发现无关特征变化导致准确率平均下降32%。
  • 提出基于大语言模型的反事实数据增强策略,使检测性能在干扰场景下平均提升56%。
  • 适合关注模型抗欺骗能力、真实场景部署的AI安全与内容审核研究者。

现有社交机器人检测器在基准测试中表现良好,但在多样化的现实场景中鲁棒性受限,主要因缺乏清晰的标签和存在多种误导性线索。特别是,模型依赖表面相关性而非因果特征的捷径学习问题未得到充分关注。为填补这一空白,我们深入研究了文本特征引发的潜在捷径对检测器的影响。通过构建用户标签与表面文本线索之间的虚假关联,设计了一系列捷径场景以评估模型鲁棒性。结果表明,无关特征分布的变化显著降低检测性能,基线模型平均相对准确率下降32%。为此,我们提出基于大语言模型的缓解策略,利用反事实数据增强,在个体用户文本、整体数据集分布及模型提取因果信息能力三个层面进行干预。该策略在捷径场景下实现平均56%的相对性能提升。

原文摘要 · Abstract (English)

While existing social bot detectors perform well on benchmarks, their robustness across diverse real-world scenarios remains limited due to unclear ground truth and varied misleading cues. In particular, the impact of shortcut learning, where models rely on spurious correlations instead of capturing causal task-relevant features, has received limited attention. To address this gap, we conduct an in-depth study to assess how detectors are influenced by potential shortcuts based on textual features, which are most susceptible to manipulation by social bots. We design a series of shortcut scenarios by constructing spurious associations between user labels and superficial textual cues to evaluate model robustness. Results show that shifts in irrelevant feature distributions significantly degrade social bot detector performance, with an average relative accuracy drop of 32\% in the baseline models. To tackle this challenge, we propose mitigation strategies based on large language models, leveraging counterfactual data augmentation. These methods mitigate the problem from data and model perspectives across three levels, including data distribution at both the individual user text and overall dataset levels, as well as the model's ability to extract causal information. Our strategies achieve an average relative performance improvement of 56\% under shortcut scenarios.

社交机器人检测大语言模型反事实增强模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。