arXiv:2505.19670cs.CLcs.MM2025-05EMNLP被引 5

通过重塑表示空间提升音频大模型安全性,避免过度拒绝有用请求。

Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models

  • 无监督微调重构模型表示空间,增强安全对齐能力。
  • 三类输入下安全性显著提升,过度拒绝率仅增加0.88%。
  • 适合关注音频模型安全与可用性平衡的研究者使用。

大型音频语言模型(LALMs)扩展了大语言模型在语音交互方面的能力。然而,最新研究发现,由于安全对齐不足,这些模型仍易受有害查询攻击。尽管文本和视觉大模型的防御技术已有进展,但针对LALMs的安全对齐策略及专用音频安全数据集仍严重缺失。基于有监督微调(SFT)的防御措施难以兼顾安全性提升与避免过度拒绝,严重影响模型帮助性。本文提出一种无监督安全微调策略,通过重塑模型表示空间,在不牺牲太多实用性的情况下增强现有LALMs的安全对齐能力。在三代Qwen LALMs上进行的实验表明,该方法在音频-文本、纯文本和纯音频三种输入模式下均显著提升安全性,同时平均仅导致0.88%的过度拒绝率上升。警告:本文包含有害示例。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions. However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safety-alignment. Despite advances in defence measures for text and vision LLMs, effective safety-alignment strategies and audio-safety dataset specifically targeting LALMs are notably absent. Meanwhile defence measures based on Supervised Fine-tuning (SFT) struggle to address safety improvement while avoiding over-rejection issues, significantly compromising helpfulness. In this work, we propose an unsupervised safety-fine-tuning strategy as remedy that reshapes model's representation space to enhance existing LALMs safety-alignment while balancing the risk of over-rejection. Our experiments, conducted across three generations of Qwen LALMs, demonstrate that our approach significantly improves LALMs safety under three modality input conditions (audio-text, text-only, and audio-only) while increasing over-rejection rate by only 0.88% on average. Warning: this paper contains harmful examples.

音频大模型安全对齐过拒问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。