arXiv:2606.21147cs.SDcs.AI2026-06被引 1

测试大模型对假有害音频的过度拒绝问题

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?

论文配图:AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?
图 1 · 摘自论文原文
  • 构建首个针对音频大模型的过拒评测基准
  • 12个模型在3000个样本上普遍存在过拒现象
  • 提出轻量级方法降低误拒,适合安全对齐研究者

大型音频语言模型(LALMs)在多种音频任务中表现优异。随着其在真实场景中的部署,确保安全性对齐愈发重要。尽管拒绝机制能防止模型回应有害请求,但也可能导致过拒——即错误拒绝无害查询。这一问题在音频领域尤为突出,因为孤立的语音可能看似有害,但在结合背景声等上下文后实际为良性。为此,我们提出AOR-Bench(音频过拒基准),首个专为LALMs设计的过拒评测基准,包含6类场景共3000个伪有害音频样本。评估12个来自6大模型系列的代表性LALMs,发现过拒现象普遍存在,并揭示了安全判断中的若干重要模式。作为初步应对,我们探索了两种轻量级策略(如思维链与激活引导)以减少过拒。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio tasks. As they are increasingly deployed in real-world applications, ensuring their safety alignment has become more important. Although refusal mechanisms serve as a key safeguard by preventing LALMs from responding to harmful requests, they can also lead to over-refusal, where models incorrectly reject benign queries. This issue is especially challenging in the audio domain because speech that appears harmful in isolation may become benign when interpreted together with the surrounding acoustic context, such as background sounds. To study this problem, we introduce AOR-Bench (Audio Over-Refusal Benchmark), the first benchmark for over-refusal specifically designed for LALMs. AOR-Bench contains 3,000 pseudo-harmful audio samples across six scenario categories. Evaluating 12 representative LALMs from six major model families, we find that over-refusal is widespread (Figure 1) and uncover several important patterns in their safety judgments. As a preliminary effort to mitigate this issue, we further explore two lightweight strategies (e.g., Chain-of-Thought and activation steering) to reduce over-refusal.

音频大模型安全对齐过拒检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。