arXiv:2510.16893cs.SDcs.AI2025-10被引 5

研究情绪变化对大音频语言模型安全性的威胁,发现中等强度情绪最危险。

Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations

  • 构建多情绪强度恶意指令数据集,测试大音频语言模型的安全性。
  • 中等情绪强度下模型生成不当回应的风险最高,非单调变化。
  • 揭示情绪波动下的安全漏洞,适合关注AI伦理与鲁棒性的研究者。

大型音频-语言模型(LALMs)将文本大模型扩展至听觉理解,为多模态应用带来新机遇。尽管其感知、推理和任务表现已广泛研究,但在副语言特征变化下的安全性对齐仍鲜有探讨。本文系统研究说话人情绪的影响,构建了包含多种情绪和强度的恶意语音指令数据集,并评估了几种前沿的LALMs。结果表明存在显著的安全性不一致:不同情绪引发不同程度的不安全响应,且强度效应呈非单调性,中等强度表达往往带来最大风险。这些发现揭示了LALMs中被忽视的安全漏洞,呼吁设计专门应对情绪变化的对齐策略,以确保在真实场景中部署时的可信性。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, and task performance have been widely studied, their safety alignment under paralinguistic variation remains underexplored. This work systematically investigates the role of speaker emotion. We construct a dataset of malicious speech instructions expressed across multiple emotions and intensities, and evaluate several state-of-the-art LALMs. Our results reveal substantial safety inconsistencies: different emotions elicit varying levels of unsafe responses, and the effect of intensity is non-monotonic, with medium expressions often posing the greatest risk. These findings highlight an overlooked vulnerability in LALMs and call for alignment strategies explicitly designed to ensure robustness under emotional variation, a prerequisite for trustworthy deployment in real-world settings.

音频语言模型安全对齐情绪识别AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。