arXiv:2607.15694eess.AScs.SD2026-07

AI克隆语音难识别,专业配音演员声纹易被误判

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

论文配图:A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
图 1 · 摘自论文原文
  • 用嵌入空间几何限制揭示声纹识别的固有误判下限
  • 即使优化算法,仍存在2.6%闭集误判率和13%会话间误判率
  • 通用模型导致半数非演员克隆误指演员,需专用训练模型缓解

声优的声音是核心资产,AI克隆直接构成威胁。传统方法通过嵌入相似度阈值识别注册声优,但在专业声优中失效:受训声音密集分布于嵌入空间,且同一演员可表现多种风格。在1,168名日本声优(56,568段音频,约63小时)上,经校准、分数归一化与判别重排序(线性/非线性,含PLDA)后,仍存在约2.6%的闭集误判下限,该误差源于嵌入几何结构,非后端模型所致。最優集成模型仍无法避免此下限;会話間誤判率经重排序降至13.0%。同样因空间拥挤,虚假指认频发:通用英文编码器下,约一半非注册者克隆会错误指向注册声优;而32%的Seed-VC克隆注册目标在相同阈值下被漏检。单一操作点无法同时避免两类错误。领域匹配、声优训练的编码器显著缓解问题(性别差距缩小四倍,误指率降至1.5%-10%),但未消除下限。控制实验(编解码、信道、声码器、内容)表明误判率反映真实-合成差异,非缺失说话人信息。因此,固定阈值克隆归属不可靠,通用编码器亦不公平。鲁棒归属需引入反欺骗感知的开放集1:N验证(反欺骗门控、领域匹配编码器、个体校准、弃权选项),即便如此,仅支持检测而非自主执行。

原文摘要 · Abstract (English)

A voice actor's voice is their asset, and AI cloning directly threatens it. The natural defense flags the enrolled actor whose embedding similarity to a suspect recording crosses a threshold. We show it fails where it is most needed: trained voices crowd the embedding space, and each actor performs many styles. On 1,168 Japanese voice actors (56,568 segments, ~63 h), a misidentification floor survives calibration, score normalization, and discriminative re-ranking (linear and nonlinear, including PLDA): the residual is a limit of the embedding geometry, not of the back-ends we evaluate. The best ensemble still leaves ~2.6% closed-set misidentification, several-fold above matched controls; session-disjoint, re-ranking lowers the floor only to 13.0%. The same crowding drives false attribution: on a generic English encoder, roughly half the clones of non-enrolled people falsely accuse an enrolled actor, while -- by a separate real-vs-synthetic shift -- 32% of Seed-VC clones of enrolled targets are missed at the same threshold; one operating point couples the two, and none escapes both. A domain-matched, voice-actor-trained encoder mitigates substantially (a four-fold gender gap vanishes; wrongful misattribution falls to 1.5-10%), but does not remove the floor. Controls (codec, channel, vocoder, content) support reading the miss rate as a real-versus-synthetic covariate shift, not missing speaker information. Fixed-threshold clone attribution is thus unreliable here, and on a generic encoder unfair. Robust attribution must extend spoofing-aware speaker verification to open-set 1:N (anti-spoofing gate, domain-matched encoder, per-speaker calibration, abstain option), and even then supports detection, not autonomous enforcement.

语音克隆声纹识别反欺骗开放集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。