提出新型语音大模型欺骗攻击,揭示其易受心理操控的弱点。
Benchmarking Gaslighting Attacks Against Speech Large Language Models
- 设计五类心理操控策略,模拟真实场景中的误导性语音输入。
- 测试显示平均准确率下降24.3%,部分模型出现自责或拒绝响应等异常行为。
- 适用于评估语音AI安全性的研究人员与产品开发者。
随着语音大语言模型(Speech LLMs)在语音应用中日益普及,确保其对操纵性或对抗性输入的鲁棒性变得至关重要。尽管已有研究关注文本大模型和视觉-语言模型的对抗攻击,但语音交互特有的认知与感知挑战仍被忽视。语音具有固有的模糊性、连续性和感知多样性,使对抗攻击更难检测。本文提出‘煤气灯攻击’(gaslighting attacks),即精心设计的提示,旨在误导、覆盖或扭曲模型推理,以评估语音大模型的脆弱性。我们构建了五种操控策略:愤怒、认知干扰、讽刺、隐含否定和专业否定,用于测试模型在多种任务下的鲁棒性。该框架不仅量化性能下降,还捕捉模型的异常行为反应,如主动道歉或拒绝执行指令。此外,通过声学扰动实验评估多模态鲁棒性。在5个语音与多模态大模型上,基于超过10,000个样本、来自5个不同数据集的综合评估显示,平均准确率下降24.3%,表明模型存在显著的行为脆弱性。这些发现凸显了构建更稳健、可信的语音智能系统的重要性。
原文摘要 · Abstract (English)
As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input becomes critical. Although prior work has studied adversarial attacks in text-based LLMs and vision-language models, the unique cognitive and perceptual challenges of speech-based interaction remain underexplored. In contrast, speech presents inherent ambiguity, continuity, and perceptual diversity, which make adversarial attacks more difficult to detect. In this paper, we introduce gaslighting attacks, strategically crafted prompts designed to mislead, override, or distort model reasoning as a means to evaluate the vulnerability of Speech LLMs. Specifically, we construct five manipulation strategies: Anger, Cognitive Disruption, Sarcasm, Implicit, and Professional Negation, designed to test model robustness across varied tasks. It is worth noting that our framework captures both performance degradation and behavioral responses, including unsolicited apologies and refusals, to diagnose different dimensions of susceptibility. Moreover, acoustic perturbation experiments are conducted to assess multi-modal robustness. To quantify model vulnerability, comprehensive evaluation across 5 Speech and multi-modal LLMs on over 10,000 test samples from 5 diverse datasets reveals an average accuracy drop of 24.3% under the five gaslighting attacks, indicating significant behavioral vulnerability. These findings highlight the need for more resilient and trustworthy speech-based AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。