arXiv:2411.08316cs.CRcs.SD2024-11

用少量目标语音合成命令,即可操控智能音箱执行敏感操作。

Evaluating Synthetic Command Attacks on Smart Voice Assistants

  • 仅需少量目标语音,通过拼接合成语音发起攻击。
  • 合成语音可让音箱执行敏感操作,成功率较高。
  • 攻击可借周边被控设备发起,隐蔽性强,适合安全研究者参考。

近期语音合成技术进步,结合大规模语音数据获取的便利性,给智能语音助手(如 Amazon Alexa、Google Home)带来新威胁。本文研究:是否可用目标用户少量无关语音,合成可被语音助手识别的指令。具体考察在语音助手将指令来源与授权用户匹配,且应用(如 Alexa Skills)仅响应经认证用户(设定置信度)指令时的攻击可行性。结果表明,即使使用简单的拼接式语音合成,攻击者也能成功操控语音助手执行敏感操作。此外,若通过附近被攻陷设备发起攻击,其主机和网络开销极小。研究揭示了对合成恶意指令防御机制的迫切需求。

原文摘要 · Abstract (English)

Recent advances in voice synthesis, coupled with the ease with which speech can be harvested for millions of people, introduce new threats to applications that are enabled by devices such as voice assistants (e.g., Amazon Alexa, Google Home etc.). We explore if unrelated and limited amount of speech from a target can be used to synthesize commands for a voice assistant like Amazon Alexa. More specifically, we investigate attacks on voice assistants with synthetic commands when they match command sources to authorized users, and applications (e.g., Alexa Skills) process commands only when their source is an authorized user with a chosen confidence level. We demonstrate that even simple concatenative speech synthesis can be used by an attacker to command voice assistants to perform sensitive operations. We also show that such attacks, when launched by exploiting compromised devices in the vicinity of voice assistants, can have relatively small host and network footprint. Our results demonstrate the need for better defenses against synthetic malicious commands that could target voice assistants.

语音合成安全攻击智能音箱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。