音频大模型需遵循最小权限原则,防范语音隐私泄露风险
The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
- 直接处理语音而非转录,保留音调、多人说话等细节
- 端到端模型易泄露说话人身份,存在偏见与情绪识别风险
- 建议采用最小权限设计,保障语音隐私安全
最新音频语言模型(Audio LMs)直接处理语音,不再依赖独立转录步骤。这一转变保留了语调、多说话人等关键信息,但也带来新安全风险,如说话人身份线索和敏感语音特征的滥用,可能引发法律问题。实验表明,相比级联流水线,端到端建模会引入身份推断、偏见决策和情绪识别等社会技术风险,令人担忧其是否存储声纹并处于现有法律框架的模糊地带。本文主张在开发与部署中采纳最小权限原则,评估隐私与安全风险,并明确信息访问范围。同时指出当前音频模型评测体系的不足,提出亟待解决的技术与政策性开放问题,以实现端到端音频大模型的负责任应用。
原文摘要 · Abstract (English)
The latest Audio Language Models (Audio LMs) process speech directly instead of relying on a separate transcription step. This shift preserves detailed information, such as intonation or the presence of multiple speakers, that would otherwise be lost in transcription. However, it also introduces new safety risks, including the potential misuse of speaker identity cues and other sensitive vocal attributes, which could have legal implications. In this paper, we urge a closer examination of how these models are built and deployed. Our experiments show that end-to-end modeling, compared with cascaded pipelines, creates socio-technical safety risks such as identity inference, biased decision-making, and emotion detection. This raises concerns about whether Audio LMs store voiceprints and function in ways that create uncertainty under existing legal regimes. We then argue that the Principle of Least Privilege should be considered to guide the development and deployment of these models. Specifically, evaluations should assess (1) the privacy and safety risks associated with end-to-end modeling; and (2) the appropriate scope of information access. Finally, we highlight related gaps in current audio LM benchmarks and identify key open research questions, both technical and policy-related, that must be addressed to enable the responsible deployment of end-to-end Audio LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。