arXiv:2509.19143cs.CLcs.AI2025-09EMNLP被引 2

自动跨语言生成对抗性提示,提升全球虚假信息检测能力

Anecdoctoring: Automated Red-Teaming Across Language and Place

  • 用多语言事实核查数据构建知识图谱,增强攻击大模型
  • 在英语、西班牙语、印地语中实现更高攻击成功率
  • 适合需要全球化对抗测试的研究者和安全团队

虚假信息是生成式人工智能滥用的主要风险之一。随着生成式AI的全球普及,亟需在多种语言和文化背景下开展稳健的红队评估(即系统性对抗探测),但现有红队数据集多集中于美国和英语地区。为此,我们提出“anecdoctoring”——一种自动跨语言、跨文化生成对抗性提示的新方法。从英文、西班牙语和印地语的三个事实核查网站及美国、印度两个地理区域收集虚假信息声明,将个体声明聚类为更广泛的叙事主题,并通过知识图谱对这些聚类进行表征,以此增强攻击型大模型。与少量示例提示相比,该方法显著提高攻击成功率,并具备更强可解释性。结果表明,虚假信息防范措施必须具备全球扩展能力,并基于真实世界中的对抗滥用场景。

原文摘要 · Abstract (English)

Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates red-teaming evaluations (i.e., systematic adversarial probing) that are robust across diverse languages and cultures, but red-teaming datasets are commonly US- and English-centric. To address this gap, we propose "anecdoctoring", a novel red-teaming approach that automatically generates adversarial prompts across languages and cultures. We collect misinformation claims from fact-checking websites in three languages (English, Spanish, and Hindi) and two geographies (US and India). We then cluster individual claims into broader narratives and characterize the resulting clusters with knowledge graphs, with which we augment an attacker LLM. Our method produces higher attack success rates and offers interpretability benefits relative to few-shot prompting. Results underscore the need for disinformation mitigations that scale globally and are grounded in real-world adversarial misuse.

红队测试跨语言虚假信息知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。