通过分析声音衰减区群延迟差异,可有效识别AI生成的瞬态声波。
Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds

- 利用群延迟在衰减区的分布差异作为检测依据
- 衰减区群延迟KL散度达0.322,显著高于起始区的0.022
- 群延迟图作输入时CNN分类准确率90%-94%,适合音频鉴伪
我们研究了是否可通过群延迟分析区分AI生成与真实瞬态声波。核心发现是:AI生成的声音在起始区域群延迟分布几乎相同,但在衰减区域表现出可测量的差异——衰减区KL散度为0.322,而起始区仅0.022。跨频带群延迟变异性的单特征AUC为0.720;基于九个衰减区特征的随机森林(RF)在样本不重叠评估下达到AUC=0.884。将群延迟图作为独立二维输入送入CNN分类器,准确率达90%–94%,表明群延迟蕴含丰富判别信息。在生成器保留测试中,CNN与Transformer分类器表现波动大(AUC 0.457–0.918)。RF在所有方法中平均保持最高持存性(66.7%),虽平均AUC(0.731)低于CNN(0.762)和AST(0.772),但未出现极端低于随机的崩溃现象。27种STFT配置下的参数敏感性分析显示,RF AUC稳定在0.700–0.847(标准差0.035)。结果表明,衰减区群延迟可作为物理可解释的音频鉴伪线索,补充幅度基分类器,但仍需更广泛验证。
原文摘要 · Abstract (English)
We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group delay analysis. Our central finding is that AI-generated impulsive sounds show near-identical onset-region group-delay distributions but exhibit measurably different group-delay behavior in the late decay region: decay-region KL divergence reaches $0.322$ compared to near-zero onset divergence ($0.022$). Cross-band GD variability achieves single-feature AUC~=~0.720, and a Random Forest (RF) over nine decay-region features reaches AUC~$=$~0.884 under sample-disjoint evaluation. A group delay map used as a standalone 2D input to CNN classifiers achieves 90--94\% accuracy, demonstrating that group delay carries substantial discriminative information. Under generator hold-out, CNN and transformer classifiers show highly variable AUC (0.457--0.918). The group delay RF achieves the highest average hold-out accuracy among the evaluated methods ($66.7\%$) and avoids extreme below-random collapse, although its average AUC (0.731) is lower than CNN avg (0.762) and AST (0.772). Parameter sensitivity analysis across 27 STFT configurations confirms that the RF AUC remains stable (0.700--0.847, std~=~0.035). These results suggest that decay-region group delay can serve as a physically interpretable forensic cue that complements magnitude-based classifiers, while broader validation remains necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。