arXiv:2601.03134cs.CL2026-01被引 2

通过模拟对话骗局,揭示大模型在多轮对抗中的行为规律。

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

  • 用双语大模型互演方式,模拟真实社交工程骗局场景。
  • 发现攻击方多轮升级策略,防御方依赖验证与延迟应对。
  • 跨模型跨语言对比显示策略响应存在系统性差异。

随着大模型在多轮对话中展现更强说服力,传统单轮安全评估难以捕捉其对抗性交互动态。本文构建受控的模型对模型仿真框架,针对中英文双语社交工程场景开展自动化红队测试。评估了八种主流大模型,分析对话结果、标注攻防策略类型,并建模双方互动机制。结果显示,多轮对抗对话呈现可重复的升级模式;防御方响应常依赖验证、延时和渠道控制。统计分析表明,不同模型与语言间的结果分布存在显著差异;过渡分析揭示防御策略对攻击手法的响应结构在语言间具有系统性变化。研究强调多轮对抗对话中交互结构的重要性,证明受控模型对模型模拟可支持对对抗对话机制的深入分析。

原文摘要 · Abstract (English)

As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in extended interaction settings that traditional single-turn safety evaluations fail to capture. We systematically study these interactional dynamics using a controlled LLM-to-LLM simulation framework for automated red-teaming across bilingual social engineering scenarios. Evaluating eight state-of-the-art models in English and Chinese, we analyze dialogue-level outcomes, annotate attacker and defender strategy families, and model interaction dynamics between them. Results show that multi-turn adversarial dialogues follow recurrent escalation patterns, while defensive responses frequently rely on verification, delay, and channel control. We further find statistically significant cross-model and cross-lingual differences in outcome distributions, and transition analysis reveals systematic structural variation in how defender strategies respond to attacker tactics across languages. These findings highlight the importance of studying interactional structure in multi-turn adversarial dialogue settings and demonstrate how controlled LLM-to-LLM simulations can support mechanistic analysis of adversarial conversational dynamics.

对话安全红队测试多轮对抗社交工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。