基于真实数据构建语音生成风险分类体系,揭示未授权语音滥用的多重风险。
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data
- 通过569起事件与2221条讨论构建多源威胁模型。
- 识别出暴露度、社会可见性等关键上下文因素如何加剧风险。
- 适合关注语音隐私与AI治理的研究者和政策制定者。
随着生成式语音模型能力提升与应用普及,未经同意收集、重用和合成语音数据正带来新型隐私、安全与治理风险,现有统一威胁模型难以充分捕捉此类风险。为填补空白,我们提出V.O.I.C.E(Voice, Ownership, Identity, Control, Expression)风险分类体系,基于多源威胁建模:涵盖569起重大AI事件数据库、联邦贸易委员会(FTC)及互联网犯罪投诉中心(IC3)的事件;1067份来自美国不同群体(包括语音演员、网络名人、政界人士及普通公众)的直接报告;以及2221条Reddit讨论。该分类体系基于真实数据,明确刻画风险产生机制,并分析其与暴露程度、社会可见性及法律保护可得性等上下文因素的交互作用。
原文摘要 · Abstract (English)
As generative voice models are rapidly advancing in both capabilities and public utilization, the unconsented collection, reuse, and synthesis of voice data are introducing new classes of privacy, security and governance risk that are poorly captured by existing, largely uniform threat models. To fill the gap, we present V.O.I.C.E, a taxonomy of voice generation risk grounded in a multi-source threat modeling effort with 569 incidents from major AI incident database, FTC and Internet Crime Complaint Center (IC3); 1067 direct incident reports from U.S. based participants across diverse groups (including voice actors, internet personalities, political personnel, and general public); and 2,221 Reddit discussions. Grounded in real-world data, our taxonomy explicitly models how risk emerges, interact with contextual factors such as degree of exposure, social visibility, and the availability of legal protections for various affected groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。