arXiv:2602.04796eess.AScs.SD2026-02中稿 · ICML被引 4

用语音语言模型评估对话安全,发现音频有文本之外的判断信息。

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

  • 构建24000条带安全风险的多轮语音对话数据集,覆盖8类风险和5级严重度
  • 音频提供文本外的非词汇证据,但多模态效果不普适且受融合机制限制
  • 适合做语音安全评估、多模态模型优化的研究者与开发者参考

语音对话中的社会不安全内容评估仍以文本为中心,忽略了语调特征和转录错误。我们提出 LALM-as-a-Judge,包含一个开放基准,涵盖24,000条多轮语音对话,每条对话中有一个局部不安全发言,源自8类社会不安全行为及5个严重度等级。我们在文本仅、音频仅和多模态三种设置下,评估6个大音频-语言模型(LALMs)作为评判者的敏感性、严重度排序准确性以及发言位置偏差。结果表明,音频提供了超越转录语义的非词汇证据;多模态增益并非普遍存在,可能表现为文本锚定、平衡、保守或干扰,这与音频路径瓶颈和融合机制局限相关。我们将该基准定位为诊断工具,并为模型、模态和提示选择提供实践指导。

原文摘要 · Abstract (English)

Evaluation of socially unsafe content in spoken dialogues remains text-centric, missing prosody and transcription failures. We present LALM-as-a-Judge, which includes an open benchmark of 24,000 multi-turn spoken dialogues with one localized unsafe turn, generated out of 8 socially unsafe categories and 5 severity levels. We evaluate 6 large audio-language models (LALMs) as judges, open and closed-source, in text-only, audio-only, and multimodal setups by their sensitivity, severity-order specificity, and turn-position bias for socially harmful content in the dialogue. Results show that audio contributes non-lexical evidence beyond transcript semantics and that multimodal gains are not universal but can be text-anchored, balanced, conservative, and interfering, which we link to the audio pathway bottlenecks and fusion limits. We position the benchmark as diagnostic and derive practitioner guidance for model, modality, and prompts choices.

语音安全多模态评估大模型评测对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。