arXiv:2604.14548cs.SDcs.LG2026-04被引 3

语音模型在多人场景下易因说话人、语调和环境出错,需新评测框架

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

论文配图:VoxSafeBench: Not Just What Is Said, but Who, How, and Where
图 1 · 摘自论文原文
  • 分两层评测:内容风险与语音情境风险并重
  • 语音线索下安全、公平、隐私防护率显著下降
  • 适合关注语音模型社会对齐的开发者与研究者

随着语音语言模型(SLMs)从个人设备转向共享多用户环境,其响应必须考虑话语之外的因素:谁在说、如何说、在哪说。这些因素可能使原本无害的请求变得不安全、不公平或侵犯隐私。现有评测大多聚焦基础音频理解,孤立分析风险,或混淆内在有害内容与由声学上下文引发的问题。本文提出VoxSafeBench,首个联合评估语音模型在安全、公平、隐私三个维度社会对齐能力的基准。采用两级设计:一级用匹配文本与音频输入评估内容风险;二级针对语音条件风险——转录文本无害,但正确回应依赖说话人、副语言特征或环境。通过中间感知探测验证,前沿模型虽能识别声学线索却仍无法恰当响应。在22项双语任务中发现,文本上表现稳健的防护机制在语音中显著退化:说话人与场景相关风险导致安全意识下降,声音传达的群体差异削弱公平性,上下文声学线索使隐私保护失效。结果揭示普遍存在的语音具身鸿沟:当前模型常识别文本中的社会规范,却无法在关键线索为语音时应用。

原文摘要 · Abstract (English)

As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than the words alone. Who is speaking, how they sound, and where the conversation takes place can each turn an otherwise benign request into one that is unsafe, unfair, or privacy-violating. Existing benchmarks, however, largely focus on basic audio comprehension, study individual risks in isolation, or conflate content that is inherently harmful with content that only becomes problematic due to its acoustic context. We introduce VoxSafeBench, among the first benchmarks to jointly evaluate social alignment in SLMs across three dimensions: safety, fairness, and privacy. VoxSafeBench adopts a Two-Tier design: Tier1 evaluates content-centric risks using matched text and audio inputs, while Tier2 targets audio-conditioned risks in which the transcript is benign but the appropriate response hinges on the speaker, paralinguistic cues, or the surrounding environment. To validate Tier2, we include intermediate perception probes and confirm that frontier SLMs can successfully detect these acoustic cues yet still fail to act on them appropriately. Across 22 tasks with bilingual coverage, we find that safeguards appearing robust on text often degrade in speech: safety awareness drops for speaker- and scene-conditioned risks, fairness erodes when demographic differences are conveyed vocally, and privacy protections falter when contextual cues arrive acoustically. Together, these results expose a pervasive speech grounding gap: current SLMs frequently recognize the relevant social norm in text but fail to apply it when the decisive cue must be grounded in speech. Code and data are publicly available at: https://amphionteam.github.io/VoxSafeBench_demopage/

语音模型社会对齐评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。